Sanitizer API vs DOMPurify for Rich Text
Define capabilities, run native and pinned-library paths over one hostile corpus, normalize the DOM, and deploy behind negative gates.
Sanitizer API vs DOMPurify is a rendering-contract decision for user-authored rich text, not permission to send arbitrary HTML into the live DOM. This comparison fixes one capability allowlist and hostile corpus, feature-detects the native safe path, pins the library path, and exports normalized removal evidence.
Sanitizer API vs DOMPurify starts with a contract
Sanitizer API vs DOMPurify is not a contest to find one magic function. The decision is a rendering contract for user-authored rich text: which elements, attributes, URL schemes, and embedding capabilities are allowed; where sanitization happens; what happens on unsupported browsers; and how the application proves that raw markup never reaches an unsafe sink.
Begin with the threat model. The input is untrusted even when it came from an authenticated customer, a staff tool, or an AI assistant. Stored content can outlive the bug or account that created it. Sanitize as close as possible to the final browser context, keep server-side validation for product rules, and apply Content Security Policy as defense in depth rather than as a replacement. The OWASP XSS Prevention Cheat Sheet reinforces context-appropriate output handling and layered defenses.
The WHATWG dynamic markup specification defines safe and unsafe DOM insertion APIs, including the evolving Sanitizer surface. Browser availability and exact signatures must be feature-detected on the clients you support.
This guide compares a native safe-API path with a pinned local DOMPurify path over one fixed benign-and-hostile corpus. It publishes normalized DOM and removal receipts, but it does not declare either route universally secure beyond its version, configuration, browser, and test corpus.
- Store and label the original as untrusted user-authored content.
- Apply one explicit element, attribute, and URL-scheme allowlist.
- Never send raw input to a live-DOM parsing sink.
- Normalize sanitized structure and export a regression receipt.
- Render only the sanitized result through a controlled component.
Reading rule: labels, values, patterns, and structure carry the conclusion; color is supplementary.
Define the rich-text capability allowlist
Write product capabilities before writing HTML sanitization options. A minimal comment editor might allow paragraphs, emphasis, strong text, code, lists, block quotes, and links. It may reject images, style, classes, forms, SVG, MathML, iframes, custom elements, event handlers, data URLs, and target attributes. Another product may need a broader set, but every capability should have an owner and abuse case.
Represent the contract as explicit allowed elements and allowed attributes by element. Add a URL policy that accepts only http, https, mailto, or same-origin relative links, then normalizes them. Decide whether links receive rel="nofollow noopener noreferrer" and whether external navigation is allowed at all.
Do not make data-* or aria-* globally available by reflex. They can expose application hooks, confuse semantics, or create styling and query selectors the renderer did not intend. When accessibility attributes are needed, allow the smallest named set and validate their values.
The generative UI safe-renderer article applies the same principle to component trees: capabilities are safer than arbitrary presentation instructions. Rich text should be a small document language, not a side door into the application DOM. The allowlist is the product specification; the sanitizer is one implementation of it.
Compare native and library ownership
A native safe insertion API can combine parsing, sanitization, and insertion in the browser implementation. That reduces application plumbing and may integrate with Trusted Types, but support and configuration details can vary as the platform evolves. A feature-detected path must have an explicit fallback; “works in my current browser” is not a deployment strategy.
DOMPurify is an application dependency with a mature cross-browser focus and a configurable policy surface. The official DOMPurify repository documents usage, supported environments, security notes, releases, and licensing. This lab pins official distribution bytes for version 3.4.16, records their SHA-256, and preserves the exact upstream dual-license text locally rather than fetching a floating CDN build.
The two Sanitizer API vs DOMPurify paths also differ in observability. DOMPurify exposes a removed-items diagnostic, but its maintainers warn that it should not be used to make security decisions. A native API may not provide the same removal detail. Build your own structural diff from a detached inert parse for testing, while keeping the sanitized output—not the diff—as the only candidate for rendering.
The native configuration uses the current sequence-shaped WHATWG contract for allowed elements and attributes. Exposure is not treated as success: if construction or configuration throws, the native path fails. Choose ownership deliberately. Native-first is reasonable when target support is proven and the fallback is equally constrained. DOMPurify-first is reasonable when consistent browser coverage and application-controlled release cadence matter. The safe contract must survive either choice.
Keep unsanitized markup out of live DOM
The central Sanitizer API vs DOMPurify invariant is simple: no unsanitized string reaches innerHTML, outerHTML, insertAdjacentHTML, document.write, a framework escape hatch, or an equivalent live-DOM parser. The lab stores hostile fixtures as JavaScript strings, displays them with textContent, and passes them only to the feature-detected native safe API on a detached element or to DOMPurify.
For native setHTML, create a detached template or div, apply the declared sanitizer configuration, then serialize the sanitized child structure. If the API or required configuration is unsupported, mark the native path unavailable. Do not silently fall back to setHTMLUnsafe or raw innerHTML.
For DOMPurify, request a sanitized DOM fragment or sanitized string under the pinned allowlist. If the product later mutates that result with unsafe libraries or copies rejected attributes back, the guarantee is gone. Treat sanitization as the last markup transformation before controlled insertion.
The sandboxing AI-generated code guide covers a different boundary. Sandboxes constrain executable documents; sanitizers constrain allowed markup. User-authored rich text should not be treated as code, and a sanitizer does not make arbitrary scripts safe to run. Keep those threat models separate.
| Capability | Native safe API | DOMPurify 3.4.16 |
|---|---|---|
| Availability | Feature detected | Locally bundled |
| Patch owner | Browser vendor and user | Application team |
| Policy | Sequence-shaped Sanitizer configuration | Configuration and hooks |
| Diagnostics | Structural diff | Removed diagnostic plus structural diff |
Reading rule: labels, values, patterns, and structure carry the conclusion; color is supplementary.
Run one benign and hostile fixture corpus
A Sanitizer API vs DOMPurify comparison needs identical input. Freeze benign fixtures for paragraphs, emphasis, code, lists, block quotes, safe relative links, safe HTTPS links, and expected entity handling. Freeze hostile fixtures for script elements, event-handler attributes, javascript URLs, data URLs, SVG handlers, iframe srcdoc, form controls, style injection, malformed nesting, and DOM-clobbering names.
Each fixture names its expected preserved text and prohibited capabilities. The test does not require byte-identical serialization between implementations because parsers can normalize case, entities, and optional structure differently. It requires equivalent allowed structure after canonical normalization and zero prohibited nodes, attributes, or schemes.
Normalize by walking the sanitized DOM: lowercase element and attribute names, sort attributes, canonicalize allowed URLs, discard parser-only wrappers, and serialize a compact tree. Compare that representation, not pretty-printed HTML. Report removed node names and attribute names from a detached inert parse as a human diagnostic.
The prompt-injection taint-tracking guide provides a useful provenance analogy: untrusted data should remain visibly tainted until it crosses a narrow, audited boundary. Sanitization changes the permitted representation, but the originating content should still be identified as user-authored in storage, logs, and moderation tools.
Interpret capability differences carefully
A capability matrix should distinguish API availability, parsing context, allowlist expression, URL hooks, custom-element policy, SVG and MathML handling, Trusted Types integration, diagnostics, update cadence, and browser coverage. A green cell means the compared version can support the declared contract, not that its defaults fit every application.
Native support may remove a JavaScript dependency but transfers part of the patch cadence to browser vendors and users. A pinned DOMPurify build gives the application a known version and consistent configuration, but the team must monitor releases and ship updates. Neither ownership model eliminates review.
Do not turn corpus survival into a benchmark leaderboard. The fixtures are regression evidence for one allowlist. Security research will find new parser interactions, browser behaviors will change, and product capabilities will expand. Record browser versions, library version, policy hash, fixture hash, and normalized output hash so a future run can explain its differences.
The CSP and Trusted Types rollout guide shows how to reduce the number of dangerous sinks and report violations. That platform policy complements Sanitizer API vs DOMPurify. It can block accidental raw insertion paths, while the sanitizer enforces the rich-text language users are allowed to author.
- Freeze benign preservation and hostile rejection fixtures.
- Run both available engines with one capability policy.
- Compare normalized allowed DOM and prohibited-capability checks.
- Block CI on any negative or preservation failure.
- Shadow without rendering, canary with CSP reports, and retain a plain-text kill switch.
Reading rule: labels, values, patterns, and structure carry the conclusion; color is supplementary.
Deploy behind differential and negative gates
Make negative tests release blockers. Every hostile fixture must lose script-capable elements, event handlers, dangerous schemes, active embeds, and product-disallowed form or style features. Every benign fixture must preserve its expected text and allowed structure. If the native and DOMPurify paths disagree, inspect the normalized tree and decide against the contract; do not choose whichever output looks nicer.
Roll out in stages. First run both paths in local and CI tests. Next shadow the candidate against redacted production samples without rendering its result, collecting only policy-safe structural metrics. Then canary one rendering path behind a reversible flag while CSP and Trusted Types reporting watch for unexpected sinks. Expand only after browser-support and moderation slices remain healthy.
Keep raw content immutable in storage when policy permits, and store sanitized derivatives with policy and sanitizer versions. That allows a tightened policy to re-sanitize originals without treating yesterday’s output as forever safe. Access to originals is sensitive and should not leak into analytics or client diagnostics.
Incident response needs a kill switch that can disable rich-text rendering, fall back to escaped plain text, invalidate sanitized caches, and identify affected policy versions. A security boundary is incomplete when it cannot be revoked quickly.
Export the sanitizer decision receipt
The local Sanitizer API vs DOMPurify lab contains a fixed benign-and-hostile corpus, an explicit minimal allowlist, feature detection with a sequence-shaped configuration for the native safe API, and exact inline DOMPurify 3.4.16 distribution bytes with their SHA-256, license text, and source provenance. It makes no runtime network request and never places the unsanitized fixtures into the live DOM.
For every fixture and engine, the lab exports engine and version, policy, availability, normalized tree, output hash, removed node names, removed attribute names, prohibited-capability checks, normalized cross-engine comparison, and pass status. When Chrome exposes the native API, every fixture must execute on both paths and produce the same allowed structure. The fixed corpus includes the official GHSA-h7mw-gpvr-xq4m forbidden-tag/add-tag regression, a DOM-fragment output-mode case, nesting, clobbering, URL, and attribute-context cases.
The visible preview uses textContent, so even a broken sanitizer would display a string rather than activate markup. DOMPurify receives hostile strings only inside its sanitizer, and the native path uses a detached element with setHTML; no fixture is inserted unsanitized into the live document. This Sanitizer API vs DOMPurify receipt is deterministic for the same browser capabilities.
The Sanitizer API vs DOMPurify receipt is not a security certification. DOMPurify removed-item data is diagnostic, native support changes over time, and a finite corpus cannot prove absence of every exploit. Revalidate sources, update the pinned build deliberately, add regression fixtures from real incidents, and test all supported browsers whenever the contract changes.
Choose the path your team can keep patched and observable, then make the rendering contract independent of that choice: explicit capabilities, no raw sinks, fixed negative fixtures, normalized differential evidence, defense-in-depth platform policy, and a plain-text escape hatch. That is a safer decision than relying on a library name or browser feature alone.
Runnable local artifact — The lab is finite regression evidence for one minimal allowlist in the current browser; it is not a security certification, universal policy, or substitute for patching, CSP, Trusted Types, moderation, and incident response.
Keep raw fixtures detached, sanitize through each available engine, normalize the allowed DOM, list removed nodes and attributes diagnostically, reject prohibited capabilities, and export deterministic JSON.