Sanitized private implementation
PHPHESKMySQLMarkdownCommonMarkGFMMermaidJavaScriptHTML PurifierGitHub ActionsMCP

Executive summary

This case study documents the modernization of the knowledge base inside a private enterprise portal built on HESK.

HESK already provided the validated help-desk foundation, authentication, categories, permissions, and article publication. Its knowledge base, however, treated HTML as the main authoring and storage format. That remained functional for browsers, but became less suitable as articles grew to include technical procedures, code blocks, tables, workflow diagrams, and collaboration with artificial-intelligence tools.

Rewriting the Portal or automatically converting the entire collection would have introduced risk without proportional benefit. The chosen solution was incremental: existing articles remain in HTML, while new articles can use Markdown as their canonical source. The Portal identifies the format of each record, applies the appropriate pipeline, and delivers sanitized HTML for reading.

The project also introduced Mermaid diagrams, an editor with server-side preview, and an integration contract that allows AI agents to retrieve, create, and append Markdown sections without manipulating presentation-oriented HTML.

The result is a hybrid knowledge base that remains compatible with its legacy content while becoming ready for collaboration among people, systems, and AI agents.

Context

The Portal evolved from HESK into a broader operational platform. Its knowledge base followed the same path: beyond short support answers, it began holding procedures, diagnostics, runbooks, and infrastructure-project documentation.

A technical article in this environment may contain:

  • heading hierarchy;
  • operational sequences;
  • checklists;
  • shell commands;
  • SQL queries;
  • decision tables;
  • cautions and closure criteria;
  • diagrams explaining states or dependencies.

HTML can represent all of these structures, but directly editing markup mixes content, presentation, and editor-specific details. It also makes automated updates more fragile because a tool must preserve tags, attributes, and layout structures that are not part of the technical message.

The requirement was not to remove HTML from the browser. It was to establish a simpler, predictable, and interoperable source before rendering.

Problem

The previous model presented four main challenges.

HTML was both source and presentation

Persisted content already contained rendering decisions. Small edits could introduce inconsistent markup, escaped entities, or differences between the editor and the published page.

Large technical documents were difficult to maintain

Code blocks, tables, and hierarchical documents required additional editing effort. The content was also less portable for review, comparison, and reuse in other tools.

Visual workflows lacked a native textual representation

Image-only diagrams are difficult to correct and maintain alongside an article. Mermaid lets the workflow definition remain versionable and editable as text.

The AI connector received a browser-oriented representation

AI models can interpret HTML, but it is not necessarily the best authoring contract. Presentation tags add noise, and partial updates can damage the document structure. Safe collaboration required the agent to know the canonical format and receive the original Markdown.

Objectives

The project was defined with explicit goals:

  1. preserve every existing HTML article;
  2. adopt Markdown only for new articles;
  3. keep both formats in the same database;
  4. render CommonMark/GFM on the server;
  5. sanitize HTML before displaying it;
  6. support Mermaid without a production CDN dependency;
  7. provide a technical editor with a trustworthy preview;
  8. guarantee visual parity between preview and published reading;
  9. support AI-assisted reading and authoring;
  10. prevent silent overwrites during concurrent edits;
  11. retain drafts and human review before publication;
  12. deliver the change through small, reversible phases.

Constraints

The legacy collection could not be invalidated

Existing articles remained useful and did not justify a mandatory migration. Any solution had to keep the previous HTML pipeline available.

The Portal would remain based on HESK

Authentication, authorization, and publication were already integrated with the broader platform. The goal was to extend one capability, not create another disconnected knowledge system.

Internal content could not rely on public services

Diagrams and rendering had to continue working when a CDN was unavailable or blocked. Internal diagram definitions also should not be sent to a third party for rendering.

Rendering could not trust its input

Markdown is not a security boundary by itself. The rendered result still required sanitization and explicit policy before reaching the browser.

AI could not publish unrestricted changes

Automated authoring needed to respect drafts, permissions, content versions, and human review.

Sanitized architecture

Hosts, addresses, credentials, internal names, and proprietary details are omitted.

Sanitized architecture of the hybrid knowledge base
01 Authors, readers, and AI agents People edit and review articles; agents retrieve content and may prepare controlled drafts or updates
02 HESK-based Portal Authentication, authorization, categories, publication state, administrative editing, and article reading
03 Shared editor and renderer Markdown authoring, server-side preview, CommonMark/GFM, Mermaid extraction, and safe page composition
04A Hybrid article table Legacy HTML is preserved while new records store Markdown as their canonical source with an explicit format
04B Knowledge-base connector Contract for retrieval, preview, draft creation, and hash-controlled updates
04C Local Mermaid assets Versioned library served by the Portal without a mandatory CDN request during reading
05 Tests and evidence CLI, browser, integration, security, migration, CI, deployment, and functional validation

The architecture preserves an essential distinction: Markdown is the authoring source, while sanitized HTML is the browser-reading representation.

Hybrid content model

Each article carries an explicit format. This lets the same navigation and reading flow handle old and new records without guessing from their contents.

FormatPersisted sourceTreatment
Legacy HTMLExisting HTMLPipeline compatible with previous articles
MarkdownOriginal Markdown textCommonMark/GFM, sanitization, and Mermaid-block preparation
Rendered HTMLDerived outputUsed for presentation, not as the preferred editing source

The project does not automatically convert HTML articles. This prevents unexpected visual changes and keeps the migration scope controlled. A legacy article only needs migration when there is an operational reason and a dedicated review.

Rendering and security pipeline

A Markdown article follows distinct stages:

  1. the Portal reads the declared format;
  2. a CommonMark/GFM renderer processes the Markdown;
  3. the intermediate HTML passes through sanitization;
  4. authorized Mermaid blocks receive controlled markup;
  5. the local Mermaid library turns diagram text into SVG;
  6. shared styles compose the final reading view.

Sanitization remains necessary because Markdown may contain links, attributes, or embedded HTML depending on parser configuration. The renderer is not treated as a complete security boundary.

The project also applies a fail-safe approach to diagrams: a Mermaid syntax error must not execute arbitrary content or compromise the rest of the article. The source text remains recoverable for correction.

Markdown editor

The administrative area gained a mode specifically for Markdown articles. The format is explicit and locked during editing, reducing the risk of accidental conversion between Markdown and HTML.

The editor provides shortcuts for:

  • headings;
  • bold text;
  • lists;
  • links;
  • fenced code blocks;
  • tables;
  • Mermaid diagrams.

The server produces the preview through the same domain responsible for final rendering. This avoids maintaining separate parsers in JavaScript and PHP.

An important issue appeared during validation: preview could look correct while the published view applied a different styling context. The fix was to share the same content class, typography rules, and component behavior between both surfaces. Preview became a representation of the real result rather than an approximation.

Local Mermaid and documentation as code

Mermaid keeps the diagram beside its procedure:

  • changes can be reviewed as text;
  • the workflow travels with the article;
  • AI can propose or correct the definition;
  • changing a node does not require editing an image;
  • the definition remains searchable.

The library is delivered as a local Portal asset. This reduces operational reliance on third parties, avoids sending diagram contents to a CDN, and pins the version executed in validation and production.

Local delivery also required build discipline. During the initial proof, the HTML and initializer loaded but the Mermaid asset was missing from the public directory. The browser displayed only the workflow source. That behavior reinforced the need to test the artifacts actually delivered by deployment, not only the source code.

Contract for AI agents

The connector declares the representation explicitly. A sanitized response for a Markdown article follows this principle:

{
  "content_format": "markdown",
  "content_markdown": "## Procedure\n\nTechnical content...",
  "content_hash": "content-version"
}

For Markdown records, the agent receives the source through content_markdown and does not need to reconstruct it from rendered HTML.

This contract improves collaboration because an agent can:

  • understand hierarchy without presentation noise;
  • preserve lists, tables, and code blocks;
  • produce Mermaid as text;
  • append a section without rewriting the entire page;
  • return content suitable for the same renderer used by the Portal;
  • identify the version used as the basis of its change.

The change does not assume that AI is unable to read HTML. It recognizes that a stable, semantic, and explicitly typed representation is a better contract for automated reading and authoring.

Concurrency and overwrite protection

The content_hash provides optimistic concurrency control.

Before updating an article, the agent works from a known version. If a person or another process changes the document in the meantime, the hash no longer matches and the operation can be rejected. The agent must retrieve the latest content before trying again.

This protects a simple but important scenario:

  1. an administrator opens and changes an article;
  2. an agent still holds the previous version;
  3. the agent attempts to append a section;
  4. the Portal detects the conflict instead of erasing the human edit.

Drafts and human governance

The connector can prepare new Markdown articles, but the workflow favors draft creation.

Human review checks:

  • technical accuracy;
  • absence of confidential data;
  • potentially destructive commands;
  • clarity of prerequisites;
  • readability of tables and diagrams;
  • correct category and audience;
  • validation and recovery criteria.

AI therefore acts as an assisted author and knowledge consumer, not an unrestricted publisher.

Phased delivery

The change was split into stages with validation between merges:

  1. local proof of CommonMark, sanitization, and Mermaid;
  2. data-model extension to distinguish formats;
  3. hybrid reading with HTML compatibility;
  4. Markdown editor and server-side preview;
  5. connector integration for retrieval and preview;
  6. draft creation and controlled updates;
  7. visual parity between editing and reading;
  8. technical documentation, runbooks, and closure.

The migration was applied by the deployment process, and each phase moved forward only after CI, delivery, and functional validation were green.

Problems found during validation

Testing incidents helped validate different layers of the system.

SymptomCause or lessonCorrection
Mermaid appeared as textLocal asset was absent from the served directoryAsset vendoring was added and verified in the build flow
Preview returned an HTTP errorWeb runtime and CLI did not share exactly the same contextBootstrap and error handling were adjusted for the real request
UI failed while reading JSONA failed response could be empty or incompleteClient began handling status and error payload defensively
Backticks became HTML entitiesSource content received presentation-oriented treatmentMarkdown began remaining intact as canonical text
Preview and article looked differentContainers and typography styles divergedRenderer context and styles were shared
Deployment failed despite valid codeRuntime dependencies were not guaranteed in the jobDependency installation and verification entered the pipeline

CLI tests demonstrated isolated libraries and contracts. The browser exposed asset, route, and JavaScript problems. Validation demonstrated integrated behavior. CI and deployment confirmed that the environment could reproduce the solution.

Results

Compatibility without forced migration

All previous HTML content remained available while new articles gained a source better suited to technical documentation.

Better authoring experience

Headings, code, tables, and workflow diagrams can be written in a readable format even before rendering.

Trustworthy preview

The administrative preview and the published page share the same typography and component behavior.

Agent-accessible knowledge

The connector provides original Markdown and format metadata, supporting structured reading and changes without fragile HTML manipulation.

Diagrams maintained with their content

Mermaid turned visual workflows into editable, searchable, and reviewable text.

Operationally validated delivery

Migrations, dependencies, tests, CI, deployment, and validation were completed before implementation closure.

No productivity percentage, time reduction, or AI-response quality claim is made because those measurements have not been prepared for public release.

Security decisions

The main controls were:

  • legacy HTML remains on its established path;
  • Markdown is rendered on the server before reading;
  • derived HTML passes through sanitization;
  • Mermaid uses controlled configuration and a local asset;
  • the browser does not receive connector secrets;
  • create and edit actions respect authentication and authorization;
  • drafts preserve human review;
  • hashes prevent silent overwrites;
  • error responses do not expose stack traces or configuration;
  • public examples omit endpoints, credentials, real identifiers, and topology.

Trade-offs

The hybrid architecture adds complexity because the Portal must maintain two reading paths and recognize which source can be edited.

That complexity was accepted because it avoids a high-risk migration and enables progressive delivery. An explicit format discriminator and tests for both paths keep the cost controlled.

The project also chose not to use a full visual Markdown editor. The interface favors predictable technical authoring. Users need to learn a few conventions, but preview and shortcuts lower that barrier.

Lessons learned

Modernization does not require erasing the legacy system

Preserving HTML while adding a new canonical format was safer than attempting to normalize the whole collection in one delivery.

Source and presentation require different responsibilities

Markdown suits authoring and exchange. HTML remains the correct browser output. Treating either one as the absolute replacement for the other would only move the problem.

Preview is trustworthy only when it shares the final pipeline

Two renderers or visual contexts tend to drift. Parity must be an architectural property rather than an occasional manual comparison.

AI-ready APIs need explicit semantics

A generic content field is insufficient when multiple formats coexist. Format, canonical source, version, and publication state belong in the contract.

Local tests do not represent the complete deployment

A CLI-validated parser does not prove asset availability, web-route bootstrap, JSON response behavior, or real page styling.

Useful AI still needs operational boundaries

Drafts, permissions, sanitization, and concurrency control improve automation without removing human accountability.

Limitations

  • older HTML articles retain the constraints of the previous model;
  • converting a legacy article requires individual review;
  • Markdown does not replace validation of technical accuracy;
  • Mermaid diagrams remain subject to syntax errors;
  • the current experience targets technical authors;
  • adoption and quality metrics still need to accumulate;
  • the Portal remains coupled to historical HESK conventions.

Next steps

The primary implementation is complete. Possible improvements include:

  • measuring Markdown adoption and post-review correction rates;
  • linting Markdown and Mermaid before publication;
  • expanding revision history and comparison;
  • creating templates by procedure type;
  • semantically indexing only authorized content;
  • improving diagram accessibility;
  • voluntarily migrating high-use HTML articles;
  • expanding negative tests for sanitization and concurrency.

Conclusion

The project began as a formatting improvement but evolved into a change in the Portal’s knowledge model.

By preserving HESK and existing articles, introducing Markdown as a canonical source, rendering Mermaid locally, and giving AI agents an explicit contract, the solution turned the knowledge base into a more structured, portable, and reusable resource.

The same content can now support human reading, technical maintenance, runbook creation, and AI-assisted collaboration without abandoning the validated operational environment.

Confidentiality statement

This study uses a sanitized architecture and examples. It contains no private Portal or connector source, credentials, addresses, hostnames, users, internal articles, real identifiers, database names, or production topology.