Migration notice. This repository was previously a one-way published export of a privately hosted project; early history contains squashed “Sync from internal source …” commits from that period.
originis nowgithub.com/maccavelli/mcplib,maintracks it directly, and normal commits/PRs land here.
Shared Go library for a fleet of MCP servers and related CLIs. It exists so those programs do not each invent stdio transport, tool hardening, LLM access, configuration, logging, or self-update.
The module path is github.com/maccavelli/mcplib. It is a library only: no
binary, no build/install target. Consumers import packages; they do not
vendor copies of the same primitives.
Requires Go 1.26.6.
The fleet (MagicTools and its sub-servers, MagicDev, Recall, Socratic Thinker, prepare-commit-msg, and siblings) used to re-implement the same MCP and LLM plumbing. The copies drifted: different provider menus, stale model catalogs, incompatible stdio shutdown, two CLI updaters, and none at all in other products.
mcplib is the ownership boundary for that shared behavior. Adding a provider, a redaction pattern, or an asset-naming rule here is how every consumer gets it. Re-implementing those pieces in an app is the defect this library exists to prevent.
Two runtime modes, one codebase.
Standalone. A process talks MCP over stdio to an IDE or agent. LLM calls
go through llmprovider with the process’s own credentials. Large payloads
stay in JSON-RPC (bounded). Self-update, when wired, talks to GitHub Releases.
Orchestrated. MagicTools spawns the same binary as a sub-server and injects environment:
| Variable | Meaning |
|---|---|
MCP_ORCHESTRATOR_OWNED=true |
This process is fleet-managed. |
MCP_LLM_ENABLED=true |
Shared LLM backplane is available. |
MCP_LLM_ADDR / MCP_LLM_TOKEN |
Backplane endpoint and bearer token. |
IsOrchestratorOwned() and NewBackplaneClient are the only supported
detectors. When the env is absent they report “not available”; callers degrade,
they do not fail the process. Individual servers must not re-bind these
variables through their own config stacks.
The rest of the design is the same in both modes:
- Canonical implementations, injected seams. The library owns data and
flow. Consumers own rendering and product lifecycle.
wizard.Prompter,selfupdate.Reporter/Confirmer/Installer, andWithSchemaDerivalexist so mcplib never imports a TUI toolkit, a CLI framework, or a service manager. - Hardened handlers.
HardenedAddToolalways recovers panics, keeps result content non-nil, and, when orchestrated, appends an__orchestrator_signalfor pipeline scoring. Schema derivation and serialized calls are opt-in. - Payload bypasses for size. JSON-RPC is the default. Under the
orchestrator,
fastpathcan write a tool result to a bounded artifact file, andhfsccan stream a large body as session log chunks, so the RPC frame never holds the payload. - Secrets stay in one place.
logging.Redactis the single redaction path for log buffers, sanitizing writers, and MCP log notifications.logging.MaskSecretis the opposite: a UI fingerprint (••••••••a75y), never for logs. - Optional dependencies stay in subpackages. Root
mcplibwraps the official MCP Go SDK.schemais the only package that importsinvopop/jsonschema.llmproviderspeaks HTTP itself — no vendor SDKs.
A typical server constructs NewMCPServer, wraps stdin/stdout with
NewStdioPipeline (128 KiB buffers, real-close on EOF, mutex-flushed
writes), registers tools through HardenedAddTool, and optionally attaches
a LogBuffer plus RegisterDiagnosticTool for get_internal_logs.
Peer-server access uses RecallClient / SocraticClient (Streamable HTTP,
circuit breaker, reconnect). Shared LLM, when orchestrated, uses
BackplaneClient (nil if the backplane is off).
| Package | Role |
|---|---|
mcplib |
Server wrapper, stdio pipeline, hardened tools/prompts/resources, orchestrator detection, backplane client, Recall/Socratic clients, diagnostic log tool, pagination, async stderr writer |
mcplib/llmprovider |
SDK-free LLM adapters, descriptors, static catalogs, live discovery, health probes |
mcplib/wizard |
Canonical “choose provider / key / model / fallbacks” flow behind Prompter |
mcplib/selfupdate |
GitHub Releases discovery, exact assets, SHA-256 integrity, locked binary replacement |
mcplib/schema |
Opt-in JSON Schema derivation for tool inputs |
mcplib/logging |
Secret redaction, sanitizing writer, UI key masking |
mcplib/fastpath |
Orchestrator artifact writes that skip JSON-RPC size limits |
mcplib/hfsc |
High-Fidelity Smart Chunking: stream extreme payloads as session logs |
NewProvider(name, apiKey, model, opts...) remains the API-key factory.
NewProviderWithSource(name, source, model, opts...) accepts a TokenSource
for OpenAI and Grok sessions. Every provider implements Generate; optional
interfaces add tools, extended thinking, and the sealed Item contract
(MessageItem, FunctionCallItem, FunctionCallOutputItem, ReasoningItem)
so callers switch on types instead of parsing vendor JSON.
A ChatGPT OAuth session uses chatgpt.com/backend-api/codex, never the
API-key endpoint at api.openai.com. A Grok OAuth session and a Grok API key
both use api.x.ai/v1; mcplib never sends either credential to the
cli-chat-proxy host or adds its private CLI header.
Claude and Gemini remain API-key-only.
Registered names: gemini, openai, claude, grok, opencode-zen,
opencode-go, huggingface, kilo, ollama. Ollama needs no API key, and
Kilo and OpenCode accept an empty one: they then send the service's anonymous
token, which free models answer and paid ones refuse with a typed error.
Gateways that speak Chat Completions share one encoder;
OpenCode still routes some models onto Responses-shaped or Anthropic/Gemini
wires.
Descriptors() is what a configuration UI reads — labels, env vars, default
base URLs, static model lists. Adding a provider is one descriptor plus a
constructor; wizards built on wizard.ConfigureLLM pick it up without a
downstream edit. ConfigureLLM never writes config and never logs a key.
It may return Kind=oauth with an empty APIKey field; OAuth session material is
saved through the caller's TokenStore, while consumers persist the returned
mode and model selection in their own schema. Orchestrated processes use
NewBackplaneClient, not this standalone credential wizard.
With Options.Discover set, ConfigureLLM lists the provider's models once
and asks for a search before each model menu. A blank search shows the
curated recommendations, as before. A query searches every usable model the
provider lists: a glob such as kilo-auto/* or *llama* matches whole ids,
and other queries match loosely (sonet, llama 8b). The same search is
offered for fallbacks. llmprovider.ListModelCatalog and
llmprovider.SearchModels expose the listing and the matcher to other callers.
Scripts that drive the wizard need one extra (blank) line before each model
and fallback selection.
A live listing is bounded at 10 seconds; Options.DiscoverLimit can shorten
that bound but not extend it. When the listing fails, the wizard's notice names
the cause, and ModelCatalog.Err carries it for other callers.
Retries are opt-in (GenerateWithRetry, GenerateItemsWithRetry,
GenerateThinkingWithRetry). A failed response is an *APIError carrying the
service's own error type and message (redacted, at most 512 bytes), or a
*RateLimitError for a plain 429. Sentinels classify it: ErrRateLimited,
ErrAuthFailure, ErrInvalidRequest, ErrProviderUnavailable, and
ErrQuotaExhausted (exhausted quota or balance; it also matches
ErrRateLimited) and ErrNotPermitted (region, data-policy, entitlement or
free-tier refusals). Every error still matches the sentinel its status matched
before. The retry helpers stop at once on a terminal APIError, on
x-should-retry: false, and on a server delay longer than 30 seconds, which
they return so the caller can reschedule. They honour Retry-After and
retry-after-ms. OAuth sessions make one forced-refresh retry after a 401.
A truncated answer is an error: a Responses answer marked incomplete, or a
tool call cut by the token limit, returns an *IncompleteError. A
length-truncated text answer keeps its text and sets Response.FinishReason.
The default HTTP client waits up to 300 seconds for response headers and 330
seconds in all, as the services' own clients do; set a context deadline for
less. Every request sends User-Agent: <name>/<version> (<os>; <arch>) mcplib/<version>. WithClientInfo(name, version) names the consuming
application, and WithSessionID(id) sets the conversation id that OpenCode
(x-opencode-session) and Kilo (X-KiloCode-TaskId) receive.
WithReasoningEffort sets the effort for every provider's thinking path.
Effort APIs send it as given. Claude 4.7 and later use adaptive thinking with
output_config.effort. Older Claude models map low to a 1,024-token budget.
Gemini maps low to thinkingLevel on Gemini 3 and to a 1,024-token budget on
Gemini 2.x. DiscoverModels on Kilo, OpenCode and Hugging Face ranks with the
provider's WithModelProfile and WithModelMetadataURL. On those gateways,
and for a ChatGPT session, it returns the curated listing without spending a
generation on a health probe.
selfupdate discovers GitHub Releases, selects the exact asset
<product>-<goos>-<goarch>[.exe], checks it against SHA256SUMS (and the
GitHub digest when present), and replaces the running binary under a lock.
It does not prove publisher signature authenticity.
The library never reads flags or calls os.Exit. The consumer binds:
- a source (GitHub),
- a version policy (strict
vMAJOR.MINOR.PATCH), - an asset selector,
- an installer (
StandaloneInstallerfor a plain binary;ManagedInstallerwhen a service must stop/start around the replace), - a reporter and confirmer.
Consumers publish through the reusable workflow
.github/workflows/publish-selfupdate-release.yml, pinned to the mcplib
module-tag commit. It accepts only a complete staged set, attests the files,
and publishes an immutable release. It never --clobbers. That workflow is
the only supported publication path for the asset contract.
make test # go test ./...
make lint # golangci-lint via .golangci.yml
make fmt vet tidyOpt-in: make vuln (govulncheck; non-zero when findings exist, including
stdlib fixes in newer Go patches) and make test-sum (gotestsum).
Architectural decisions live in docs/. After changing this
module, re-test the consumers that import the packages you touched.