Buyer-visible privacy gap
Protected main exposes python -m pg_llm_batch count-tokens --model <model> --text <text>. build_parser() requires --text, and _dispatch() passes args.text directly to TokenCounter.count_tokens(). Python command-line arguments are represented through sys.argv, so a prompt supplied this way becomes process-command-line material before pg-llm-batch can apply any package-side privacy control. Prompts may legitimately contain PII, confidential business text, unreleased source, incident data, or other purpose-bound content that must remain usable rather than blanket-masked.
This is distinct from #85, which removes provider secret input from argv. Prompt/content input is not necessarily a credential, but exposing it through shell history/process inspection is still an avoidable confidentiality boundary. MITRE CWE-214 explicitly covers processes invoked with sensitive command-line arguments or other visible process elements.
Bounded target
Provide a content-preserving input path for token counting that does not require prompt text in argv, while retaining the current deterministic PostgreSQL pg_tiktoken authority.
Acceptance criteria:
- add a realistic RED CLI regression showing the current canonical invocation requires prompt content in parsed command-line arguments;
- support a bounded stdin and/or reviewed file-descriptor/file input mode for
count-tokens so sensitive prompt bytes need not be present in argv or shell history;
- make source selection explicit and unambiguous: exactly one content source is accepted, conflicting sources fail closed, and omission fails with a bounded actionable error;
- bound accepted input size before full materialization or tokenization according to a documented operator/resource policy; do not introduce an unbounded stdin read merely to remove argv exposure;
- preserve prompt bytes/text needed for the actual token-counting job rather than masking or semantically rewriting content;
- define UTF-8/newline semantics explicitly so stdin/file and API token counts cannot silently disagree because of transport framing;
- diagnostics, exceptions, telemetry, and parser errors must not echo rejected prompt content;
- preserve model validation/tokenizer authority and the
pg_tiktoken-backed result; no Python-side tokenizer substitute;
- update README/API/CLI/operability/privacy guidance so argv prompt input is not presented as the production procedure;
- prove Python 3.10/3.12/3.14, exact 100% owned production statement/branch coverage and public docstrings, package/security/SAST, and current exact-source gates on the final source.
Dependency / writer boundary
Do not implement this as a competing pg_llm_batch/cli.py edit while #85 owns CLI secret-input semantics and #87 owns adjacent CLI/resource-lifetime composition. Re-evaluate from their protected integrated results or reviewed successors, then add the smallest compatible content-input boundary. This issue must not be solved by weakening #85's secret handling, by caching prompt plaintext indefinitely, or by blanket PII masking.
Primary references
MITRE. (2026). CWE-214: Invocation of process using visible sensitive information (CWE Version 4.20). https://cwe.mitre.org/data/definitions/214.html
Python Software Foundation. (2026). Command line and environment — Python 3.14.6 documentation. https://docs.python.org/3.14/using/cmdline.html
Python Software Foundation. (2026). sys — System-specific parameters and functions — Python 3.14 documentation. https://docs.python.org/3.14/library/sys.html
Buyer-visible privacy gap
Protected
mainexposespython -m pg_llm_batch count-tokens --model <model> --text <text>.build_parser()requires--text, and_dispatch()passesargs.textdirectly toTokenCounter.count_tokens(). Python command-line arguments are represented throughsys.argv, so a prompt supplied this way becomes process-command-line material before pg-llm-batch can apply any package-side privacy control. Prompts may legitimately contain PII, confidential business text, unreleased source, incident data, or other purpose-bound content that must remain usable rather than blanket-masked.This is distinct from #85, which removes provider secret input from argv. Prompt/content input is not necessarily a credential, but exposing it through shell history/process inspection is still an avoidable confidentiality boundary. MITRE CWE-214 explicitly covers processes invoked with sensitive command-line arguments or other visible process elements.
Bounded target
Provide a content-preserving input path for token counting that does not require prompt text in argv, while retaining the current deterministic PostgreSQL
pg_tiktokenauthority.Acceptance criteria:
count-tokensso sensitive prompt bytes need not be present in argv or shell history;pg_tiktoken-backed result; no Python-side tokenizer substitute;Dependency / writer boundary
Do not implement this as a competing
pg_llm_batch/cli.pyedit while #85 owns CLI secret-input semantics and #87 owns adjacent CLI/resource-lifetime composition. Re-evaluate from their protected integrated results or reviewed successors, then add the smallest compatible content-input boundary. This issue must not be solved by weakening #85's secret handling, by caching prompt plaintext indefinitely, or by blanket PII masking.Primary references
MITRE. (2026). CWE-214: Invocation of process using visible sensitive information (CWE Version 4.20). https://cwe.mitre.org/data/definitions/214.html
Python Software Foundation. (2026). Command line and environment — Python 3.14.6 documentation. https://docs.python.org/3.14/using/cmdline.html
Python Software Foundation. (2026). sys — System-specific parameters and functions — Python 3.14 documentation. https://docs.python.org/3.14/library/sys.html