How to Test a Nondeterministic LLM Application
When an application calls an LLM, the same input can produce a different response on the next run. A test that compares…
A password or access token that ends up in one exception message does not stay in one place. Logs go to local files, container output and crash reporters, and from there into traces, search indexes, archives and support exports. One leaked value spreads far beyond the service that wrote it.
Log redaction removes, replaces or transforms sensitive values before people or systems can read them. The safest approach is to keep sensitive values out of logs entirely. Log only what diagnosis needs, apply field rules before a log event is serialized, and treat masking in later layers as a backup.
Start with an inventory of every value that can end up in a log event. Look beyond your own logger calls. Request logging, database drivers, exception handlers, debug statements, HTTP clients and third-party packages all write log records.
The OWASP Logging Cheat Sheet names passwords, session identifiers, access tokens, database connection strings, encryption keys and payment card data as values that should usually not be recorded directly. They should be removed, masked, sanitized, hashed or encrypted instead. Personal data and operational details need a judgment call. An internal hostname is harmless in one environment and reveals your infrastructure in another.
A useful inventory covers these groups:
Authorization headersCombinations matter as well. An email address alone reveals little. Together with an IP address, an account number, timestamps and a transaction history, it identifies a person and what they did. There is also no standard list to copy. A study of privacy in software logs found inconsistent definitions of what counts as sensitive and no standard guidelines for anonymization. Each application needs its own policy, based on its data, what investigations need and its legal obligations.
Keeping a value out of the log is stronger than cleaning it up later, because a value that never enters a log cannot leak through another telemetry path. Replace whole-object logging with an allowlist of fields that help diagnosis.
Bad:
User object: {
name,
email,
address,
passwordHash,
accessToken,
...
}
Better:
event=user_profile_update
user_reference=[TOKENIZED]
fields_changed=["phone_number"]
result=success
trace_id=4bf92f3577b34da6a3ce929d0e0e4736
The second event records the outcome, a tokenized user reference, the name of the changed field and a trace ID without copying the user object. Do the same for HTTP traffic and SQL. Log the route and status code instead of the full body, a sanitized path instead of a query string that might contain a token, and the statement with placeholders instead of literal values.
Give each field a sensitivity class and a permitted treatment:
| Class | Examples | Default treatment |
|---|---|---|
| Public | Event name, status, currency | Keep |
| Internal | Service name, trace ID, normalized route | Keep if needed |
| Confidential | Email, IP address, order details | Minimize, partially mask or tokenize |
| Regulated | Health data, government ID, payment data | Omit or tightly control |
| Secret | Password, token, private key, session cookie | Never log; replace with a fixed marker |
Use consistent field names such as password, authorization, access_token, cookie, card_number and ssn, so that application rules, collector filters and tests all match the same fields. Treat spelling variants such as apiKey, api_key and x-api-key as the same field.
Redaction should happen before a log event is serialized, written to disk, printed to standard output or sent to a telemetry agent. Once a raw value reaches a local file, a crash dump or a remote collector, later controls can miss some of the copies.
Structured logging stores named fields instead of one block of text. A logger can then apply a rule to the authorization field without guessing whether a string contains a credential. OpenTelemetry's log model is built on structured records with typed fields, which fits this approach.
Define a schema for each important event type: the fields it may contain, their types and how each field may be transformed. Keep correlation IDs apart from values that identify a user. A trace ID connects events across services without standing in for an email address or session token. Make the safe path the default, for example with a sanitized request logger, so that services do not serialize whole request objects.
Pick the treatment that reveals no more than the diagnostic purpose needs:
[REDACTED] shows that the field existed without keeping its value.A stable pseudonym still links events. Anyone who sees the same token in several events can connect them, even without recovering the original value. A plain hash is weak for guessable values such as email addresses, phone numbers or IP addresses, because an attacker can hash likely values and compare the results. That is why the pseudonym should be keyed.
Field rules are the main control. Pattern matching adds coverage for secrets inside free text, such as JSON values, JWTs, private-key markers, card numbers, connection strings and authorization headers.
Pattern-only scrubbing has known gaps. It misses URL-encoded or Base64-encoded content, values split across fields, nested objects, malformed tokens, multiline exceptions and custom credential formats, and it mistakes ordinary numbers for card data. Use field names, deterministic patterns and specialized detectors together, and test them against the data your application produces.
The Log4j documentation makes the same point. It strongly discourages masking with its rewrite appender, because passwords and card numbers appear in too many formats to detect them all, and says the better approach is not to log such data at all. It treats regex-based masking only as defense in depth.
Error and debug paths leak most often: exception messages and stack traces, debug and trace modes, raw print statements, and database driver errors that contain query parameters. Redact exception context before serialization. When an error message can contain request data, log a stable error code and the trace ID instead. Make sure production debug settings cannot bypass the configured logger.
Application code is the earliest place to redact. Collectors and managed logging platforms catch some of what gets through, and they cover legacy applications, third-party output and services written in other languages.
Put the common filters for headers, cookies, request metadata, structured attributes and exception context in shared middleware. Run it before JSON serialization and before the logger hands the event to a file writer or exporter. Shared middleware helps several services follow the same rules, but check its coverage. A library that writes straight to a file, raw standard output, other logger configurations and startup failures can all bypass it.
Agents and collectors can apply the same field and pattern rules across all services. That helps with legacy services you cannot change yet and with third-party telemetry you do not control. The OpenTelemetry .NET redaction guide shows the SDK version of this, a processor that rewrites log attributes before export, including regular-expression rules for card numbers and email addresses. It also points to Collector processors for the same job.
Collector masking is still a late control. The original event may already sit in a local file, a container runtime buffer, a crash dump, a queue or another telemetry destination. A failed agent, a bypassed pipeline or an unrecognized field name also sends raw data onward. Keep the application-level controls even when the collector has strong detectors.
Managed log platforms can detect and mask credentials, payment data, personal data and health data when logs arrive. Amazon CloudWatch Logs, for example, supports managed and custom data identifiers, and only users with the logs:Unmask permission see the unmasked values. It masks data when it is ingested, so events that arrived before you set the policy stay unmasked.
A platform policy also does nothing for exported files, archived indexes, backups, support tickets, screenshots or data in another account. And masking is not deletion. A privileged user or system can still recover the original value.
Protect logs even after redaction:
The OWASP logging guidance recommends restricting and reviewing access to logs. Microsoft's Azure Monitor security guidance recommends filtering or obfuscating sensitive data and exporting audit data to immutable storage when records must be tamper-proof.
Redaction rules break when schemas, frameworks, vendors and telemetry destinations change, so test them like any other security control.
Unit tests should push sensitive values through your serializers, logger wrappers and filters. Cover nested objects, nulls, empty strings, Unicode, escaped characters, multiline values, URL-encoded and Base64-like strings, headers, cookies, exception messages and large request bodies.
Integration tests should look at every place that people or attackers can read:
A test that only checks the final search index misses a leak into a local file or a crash reporter.
Make builds fail when synthetic secrets, private-key delimiters, unauthorized card values, authorization headers or sensitive query parameters show up in test logs. Use realistic field names and nested structures so the tests go through the same code paths as production.
In production, count redaction events and alert on unusual increases. Record the service, the field name and the matching rule, never the original secret. Review false positives so nobody disables a useful detector, and look for false negatives through bypass reports and periodic scans.
Assume that a secret that reached a log destination has been copied. Revoke or rotate the credential first. Then find the affected indexes, files, exports, archives, backups and support artifacts, restrict access while you investigate, review the access records, and purge the data where that is technically and legally possible.
Next, fix the source and add a regression test that reproduces the leak. Deleting one event from the main index does not prove that every copy is gone. Regulated data can also trigger notification, retention or incident-reporting duties, depending on the jurisdiction and the type of data.
Give Vroni a GitHub issue, bug report, spec, or rough idea. It reads the repo, plans the change, writes code, runs checks, and works toward a review-ready pull request.
Take a look at vroni.com