Skip to content

[Privacy] Keep count-tokens prompt content out of process argv #116

Description

@seonghobae

Buyer-visible privacy gap

Protected main exposes python -m pg_llm_batch count-tokens --model <model> --text <text>. build_parser() requires --text, and _dispatch() passes args.text directly to TokenCounter.count_tokens(). Python command-line arguments are represented through sys.argv, so a prompt supplied this way becomes process-command-line material before pg-llm-batch can apply any package-side privacy control. Prompts may legitimately contain PII, confidential business text, unreleased source, incident data, or other purpose-bound content that must remain usable rather than blanket-masked.

This is distinct from #85, which removes provider secret input from argv. Prompt/content input is not necessarily a credential, but exposing it through shell history/process inspection is still an avoidable confidentiality boundary. MITRE CWE-214 explicitly covers processes invoked with sensitive command-line arguments or other visible process elements.

Bounded target

Provide a content-preserving input path for token counting that does not require prompt text in argv, while retaining the current deterministic PostgreSQL pg_tiktoken authority.

Acceptance criteria:

  • add a realistic RED CLI regression showing the current canonical invocation requires prompt content in parsed command-line arguments;
  • support a bounded stdin and/or reviewed file-descriptor/file input mode for count-tokens so sensitive prompt bytes need not be present in argv or shell history;
  • make source selection explicit and unambiguous: exactly one content source is accepted, conflicting sources fail closed, and omission fails with a bounded actionable error;
  • bound accepted input size before full materialization or tokenization according to a documented operator/resource policy; do not introduce an unbounded stdin read merely to remove argv exposure;
  • preserve prompt bytes/text needed for the actual token-counting job rather than masking or semantically rewriting content;
  • define UTF-8/newline semantics explicitly so stdin/file and API token counts cannot silently disagree because of transport framing;
  • diagnostics, exceptions, telemetry, and parser errors must not echo rejected prompt content;
  • preserve model validation/tokenizer authority and the pg_tiktoken-backed result; no Python-side tokenizer substitute;
  • update README/API/CLI/operability/privacy guidance so argv prompt input is not presented as the production procedure;
  • prove Python 3.10/3.12/3.14, exact 100% owned production statement/branch coverage and public docstrings, package/security/SAST, and current exact-source gates on the final source.

Dependency / writer boundary

Do not implement this as a competing pg_llm_batch/cli.py edit while #85 owns CLI secret-input semantics and #87 owns adjacent CLI/resource-lifetime composition. Re-evaluate from their protected integrated results or reviewed successors, then add the smallest compatible content-input boundary. This issue must not be solved by weakening #85's secret handling, by caching prompt plaintext indefinitely, or by blanket PII masking.

Primary references

MITRE. (2026). CWE-214: Invocation of process using visible sensitive information (CWE Version 4.20). https://cwe.mitre.org/data/definitions/214.html

Python Software Foundation. (2026). Command line and environment — Python 3.14.6 documentation. https://docs.python.org/3.14/using/cmdline.html

Python Software Foundation. (2026). sys — System-specific parameters and functions — Python 3.14 documentation. https://docs.python.org/3.14/library/sys.html

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions