Open source · ROSH™ Company Labs

llm-stream-guardrails redact PII and secrets from an LLM stream

Boundary-safe streaming redaction for Node and edge runtimes. It takes personal data, API keys and secrets out of a language model's reply while it is still streaming — without buffering the whole response, and with output byte-identical to filtering the finished text. TypeScript, MIT, zero dependencies.

Free and MIT-licensed. If it saves you an afternoon:
MIT licence Zero dependencies TypeScript, ESM, Node 18+ Published with a provenance attestation

Why streaming bypasses most guardrails

A model sends its answer in small pieces. Scan each piece on its own, and anything that straddles two of them goes through untouched:

chunk 1:  "...your card is 4111 11"
chunk 2:  "11 1111 1111, charge it."

Neither chunk holds a card number. Together they do — and by the time the full reply exists, the reader has already seen it. The usual workaround is to stop streaming: buffer the whole answer, filter it, then send it. That gives up the only reason to stream in the first place.

PII split across SSE chunks: a known gap, not a niche one

The same failure keeps being reported across the ecosystem:

  • LiteLLM #41611, 17 September 2026 — “Streaming guardrails: value split across two SSE chunks can pass per-chunk checks”.
  • Mastra #23783, 13 September 2026 — “PIIDetector emits PII in the clear when a match is split across stream chunks”.
  • LangChain #35011, 4 February 2026 — “Streaming bypasses guardrails/middleware”.

Until September 2026 the Vercel AI SDK said it in the guardrail example itself: “streaming guardrails are difficult to implement, because you do not know the full content of the stream until it's finished”. That comment is gone because the example was implemented — by buffering each text block until it ends, which the docs note delays output and uses memory proportional to the block size.

The guarantee: byte-identical output at any chunk size

For any input, any policy and any chunk size, the streamed output is byte-identical to filtering the whole string at once — and every chunk that leaves the library is independently well-formed.

It is enforced in CI by a property test across input, policy and chunk size, rather than promised in a README. The demo is the short version of the same argument: drag the chunk size down to a single character and the output does not change.

Open the live demo — it runs the published build in your browser. Nothing you type leaves your device.

Install — TypeScript, ESM, Node 18+

npm install llm-stream-guardrails
import { sieve } from 'llm-stream-guardrails';

for await (const token of sieve(result.textStream)) {
  process.stdout.write(token);   // secrets already gone
}

ESM only, Node 18 or newer. Any async iterable of strings goes in: the Vercel AI SDK's textStream directly, and the OpenAI and Anthropic SDKs with a one-line map to pull the text out of their event objects. For Web Streams and edge runtimes there is createSieveTransform(), a TransformStream.

What it detects: PII, API keys and secrets

  • Email addresses, phone numbers and US social security numbers.
  • Credit cards checked with Luhn, and IBANs checked with mod-97 — so a 16-digit order reference is left alone.
  • 35 credential shapes, among them OpenAI, Anthropic, Stripe, GitHub, GitLab, AWS, Google, Slack, SendGrid and npm tokens, JWTs, and complete PEM private key blocks.
  • Field names as well as values, because a model writes records as labelled lines.
  • IP addresses, opt-in, since they appear constantly in technical output.

Detection runs against a canonical view of the text, so a fullwidth digit, a non-breaking space, a zero-width character or a homoglyph cannot carry a value past a pattern written in ASCII.

What it is not

This is defence in depth. It is not a compliance control.

It matches formats, not meaning: a secret spelled out in prose, or a credential shape invented last month, gets through. Measured recall against third-party corpora that nobody here wrote is 47% to 81%, not 100%. It is not certified against GDPR, HIPAA, PCI-DSS or any other regime, and no such claim is made here — those are properties of a system and its operator, never of a dependency. Put it in front of a stream that should not carry secrets, and keep doing everything else you already do.

Numbers

  • 79 tests, including a red-team suite and a property test that compares streamed output against whole-string filtering at every chunk boundary.
  • No false positives on a 39-sample clean corpus — code identifiers, prices, order numbers and deliberate near-misses that must come through untouched.
  • Version 0.7.6, published with a verified provenance attestation, so the tarball can be traced back to the workflow that built it.
  • First week on npm (21–22 September 2026): 353 downloads, and 268 clones of the repository from 95 unique cloners.

Upstream: a documentation fix merged into the Vercel AI SDK

While this was being written, two problems turned up in the Vercel AI SDK's Language Model Middleware documentation: the streaming guardrail example left wrapStream unimplemented, and all five middleware examples omitted a required field. Issue #21209 reported both.

Three pull requests referencing that issue were merged on 22 September 2026 — #21210 on main, #21211 on release-v6.0 and #21212 on release-v5.0. They were opened by the project's own automation, and the merged commit on the default branch carries Co-authored-by: roshcompanylabs.

The defect class has a name now

On 22 September 2026 Ninad Phalak published Split-Boundary Leaks in Streaming Guardrails (Zenodo, CC BY 4.0, doi.org/10.5281/zenodo.22909585): a report collecting four independent implementations of this same defect — LiteLLM, Mastra, LangChain and NVIDIA NeMo Guardrails — alongside the unimplemented example in the Vercel AI SDK docs. It names the invariant all four lack: “For any way of splitting the same input, the concatenated streamed output must be byte-identical to filtering the whole string at once.”

That is the guarantee at the top of this page, and the design the report calls the only one that preserves streaming while satisfying the invariant — computing a settlement point per chunk — is the design this library implements. The report also records where the Vercel documentation fix came from, crediting this account by name: an outside maintainer of a streaming redaction library who reported, on 20 September 2026, that the middleware examples did not compile and that the streaming guardrail example was left unimplemented.

Two of its open questions are worth repeating, because this project answers one of them. The report notes that no project publishes chunking invariance as a stated contract, and that the latency cost of the settlement design is unmeasured across implementations. The invariant is this library’s stated contract, and the hold-back figures in the FAQ come from the benchmark that ships in the package. The report does not review this library and makes no claim about it — it is cited here because it named the class. We wrote the class up in full — the four projects, the invariant and the test that catches it — in PII split across stream chunks.

Maintainer

Maintained by Redouane — ROSH™ Company Labs. Bugs and feature requests belong in GitHub issues. A way past a detector is the most useful report this project can get; SECURITY.md explains how to send one privately.

FAQ

Questions developers ask about this

My guardrail checks every chunk and a full email still went through. How?

Each chunk is checked on its own, so a value split as "user@exa" + "mple.com" matches nothing in either half — and the browser reassembles it on screen. The check has to run on the accumulated text, not on each delta.

My middleware does detect it, but the user has already seen the text. Why does redacting not help?

A stream is append-only for the reader. Once a token is emitted you cannot take it back, so a verdict that arrives after the tokens changes your logs, not the screen. The decision has to happen before the bytes leave.

Is there a way to scan the content first and still stream it?

Yes, if you hold back only the part that could still become a match. This library emits every byte it has finished scanning and withholds the unresolved tail, so output starts immediately instead of after the full generation.

Does it buffer the whole response? What does it cost in latency?

No. On text written with spaces it holds back about twenty characters: median 18 and mean 18 over 1,792 push() calls across four corpora of spaced text and eight chunk sizes, re-measured on 0.7.6 — the benchmark ships with the package as npm run latency. On Japanese, Chinese and Thai it holds almost nothing: a 330-character Japanese reply held zero and emitted from the first character, because an ideograph cannot be part of any credential the detectors match. Near a value that might still be growing it holds more, by design, and there is no buffer-size setting, so no configuration can make it leak.

Does it stream Japanese, Chinese, Korean or Thai?

Yes, since 0.7.3 — and not before it. In 0.7.2 the engine treated a space-less paragraph as one unbroken token and released the whole block only at the end: a 2,640-character Japanese reply emitted nothing until it finished. Measured here on 0.7.3, the same kind of reply holds zero characters and emits from the first one, and an email, a card number and an API key embedded in Japanese are still redacted identically to filtering the whole string, at every chunk size tested. If you are pinned to 0.7.2, upgrade.

What exactly is the guarantee?

For any input, any policy and any chunking, the concatenated streamed output is byte-identical to filtering the whole string at once, and every emitted chunk is independently well-formed. It is a property of the engine, not a best-effort claim.

How is that tested?

A property test replays inputs across policies and chunk sizes, including single-character chunks and splits placed inside a match, and asserts the joined output equals the whole-string result. There are 79 tests in total, including a red-team suite.

Is this just holding back N bytes of the tail?

No. A fixed hold-back fails whenever a match is longer than N, and it fails silently — the output still looks redacted. Here the settlement point is derived from the patterns themselves: the furthest position no match could still grow past.

Does it work with the Vercel AI SDK, LangChain, or in a proxy?

It is a stream transform, not a framework integration, so it sits wherever you already hold the chunks: inside wrapStream, in a proxy handler, or around any async iterable of strings. createSieveTransform() gives you a TransformStream for Web Streams and edge runtimes.

Can I change what it detects and what happens on a match?

Yes. The action is mask, block or report; you can replace the detector set, add banned words, pass your own mask function, and take an onDetect callback that fires exactly once per detection with the complete value.

Can Unicode tricks get a value past it?

Detection runs on a canonical view of the text, so fullwidth digits, non-breaking spaces, zero-width characters and homoglyphs are folded before matching. Normalisation can be turned off, and doing so makes the detectors trivially evadable.

What does it not catch?

It matches formats, not meaning, so a secret written out in prose or a credential shape invented last month gets through, and a pattern glued to word characters is left alone on purpose. Measured recall on third-party corpora is 47% to 81%.

Is this a compliance control for GDPR, HIPAA or PCI-DSS?

No, and no such claim is made. It is one layer on the output path — defence in depth. Compliance is a property of a system and its operator, never of a dependency.

Does it break tool calls if the arguments are split across chunks?

It filters the stream you hand it, so give it the text channel. Tool-call JSON fragments across chunks exactly as prose does, and masking inside arguments would corrupt the call — filter that channel separately, or leave it alone.

Does it keep a sliding window across chunks, or check each chunk on its own?

It keeps the accumulated text and recomputes a settlement point on every chunk: the furthest position no match could still grow past. There is no fixed window, and a match touching the end of the buffer is never treated as final.

The AI SDK wrapStream guardrail example was left empty. What goes in it?

Since the September 2026 documentation fix it carries a working example: it buffers each text block until text-end and redacts the complete text, which closes the split-match hole — and, as the docs note, delays output and uses memory proportional to the block size. To keep the stream flowing instead, map the text deltas through sieve() inside wrapStream and re-emit them.

Why not call the non-streaming endpoint and replay the answer as fake tokens?

That works, and some maintainers suggest it, but it pushes time-to-first-token out to the full generation time — the thing streaming exists to avoid. Holding back only the unresolved tail keeps a stream a stream.

What has been fixed since launch?

Four things, each in the changelog. 0.7.3: text in a language written without spaces was held to the end instead of streaming. 0.7.4: the guard that avoids cutting between a letter and its mark knew only Latin combining marks, so Arabic, Hebrew, Devanagari and Thai marks were not treated as marks. 0.7.5: one pattern was unbounded, so scanning 8,000 unbroken characters took 846 ms where ordinary prose took 3 ms — measured here at 18 ms after the fix — and a JWT with a payload over 1,024 characters was not detected at all. 0.7.6: a caller who lowered maxRetention could ask for a ceiling shorter than the longest possible match; that configuration is now unreachable. None of the four was caught by the test suite before release. Each came from a question or a measurement afterwards, and each is covered by a test now.

How mature is it?

It was first published on 21 September 2026, and 0.7.3 followed a day later to fix streaming in languages written without spaces. The guarantee is stated precisely and covered by 79 tests you can read, but it has not been run at scale by anyone other than its author. Read the test suite before you trust it, and open an issue if you get past a detector.