An open agent-model family built for experiential intelligence: learning from experience accumulated in a real environment, and continuing to learn after deployment.
Try API | Run locally | HF Models | Technical report | Evaluation
Macaron-V1 ships two releases that share one architecture. Both keep a large base frozen, layer four LoRA specialists on top of it (Chat, Agent, Coding, GenUI), and select exactly one specialist per user turn through a shared runtime. The specialists are portable: new ones can be registered against the same base without retraining it, which is the substrate for continual learning and for composing adapters trained by different teams.
Venti is the flagship. Tall is the same four-specialist design on a smaller
base for local and lower-latency deployment. Neither is a fine-tune of the
other; they are separate bases running the same Mixture-of-LoRA (MoL) design.
| Macaron-V1-Venti | Macaron-V1-Tall | |
|---|---|---|
| Release label | 748B | 50B |
| Composition | 744B frozen GLM-5.2 base + 4 LoRA specialists | 35B-A3B frozen Qwen3.6 base + 4 LoRA specialists |
| Adapter config | rank 16, alpha 32, 7,688,042,496 stored values each | rank 64, alpha 128, 3,775,651,840 stored values each (L2 stored in F32) |
| Context | 1M on B300-class configurations; 262K in the reference vLLM profile | 262K |
| Best fit | Frontier-quality personal agent, long-horizon tool use, GenUI, SWE and terminal work | Local or on-prem deployment, latency-sensitive serving, lower cost per request |
| Deployment | 8 GPU (H20 or B300 class) via vLLM or SGLang, plus the MoL proxy | 8 GPU reference profile via vLLM, plus the MoL proxy |
| Cost profile | One resident base serves all four specialists; ~774.8B logical parameters resident against 2.976T for four separately merged base copies | Same sharing property at roughly 1/15 the base footprint |
| Weights | mindlab-research/Macaron-V1-Venti | mindlab-research/Macaron-V1-Tall |
A third build, Macaron-V1-Coding-Venti, merges the coding LoRA directly into the base instead of serving it as an adapter. It is excluded from every evaluation claim below.
Release labels are release-facing, not residency counts. Venti's four public adapter headers sum to about 30.8B stored values, so the stored total is approximately 774.8B rather than 748B. Tall's Hugging Face metadata reports 36B because it counts the base checkpoint alone; the 50B label counts base plus adapters, which is the accounting basis that makes it comparable to Venti.
We are releasing the model weights and the technical report for Macaron-V1, a 748B open-source agentic model built for personal intelligence and continual learning. Alongside the weights we are releasing the methodology and infrastructure behind it: the AutoResearch platform, Model-Harness Co-design, and the Macaron ChatBench and LivingBench suites.
Three key points:
- Releasing the model weights and technical report of Macaron V1.
- Macaron V1 is our deliberate experiment: a 748B open-source agentic model designed for personal intelligence and continual learning.
- A novel model architecture, Mixture of LoRA, routes requests across four domain experts over a frozen base, not just fine-tuning.
- Alongside Macaron V1, we are releasing the methodology and infrastructure behind it: the infrastructure and platform of AutoResearch; Model-Harness Co-design; Macaron ChatBench and Macaron LivingBench.
The hosted API is OpenAI-compatible.
| Endpoint | Base URL |
|---|---|
| China Mainland | https://mintcn.macaron.xin/ |
| Internal trial | https://mint.macaron.im/ |
To serve the weights yourself, use the
Mixture-of-LoRA-Harness,
which supplies the L0 router prompt, per-adapter conversation views, and
same-request adapter switching in front of a vLLM or SGLang engine. The harness
requires Linux with CUDA, Python 3.10+, eight visible GPUs for the checked-in
profiles, and either vLLM 0.24.0 or SGLang 0.5.13.post1 installed separately
against your host CUDA stack. Weights are not bundled with the harness; pull
them from Hugging Face and place the base shards and loras/L0-L3 under the
model root the launch scripts expect.
Validate a profile before starting an engine:
VLLM_CONFIG_CHECK_ONLY=1 ./mol_harness/scripts/start_vllm_venti.shThen bring up the engine and proxy together. The proxy listens on 8200 and the engine on 127.0.0.1:8000:
MOL_API_KEY_FILE=/path/to/api-key ./mol_harness/scripts/restart_vllm_tp8_venti.shSwap venti for tall to run the smaller profile. Both profiles target the
same GPUs and ports, so only one runs at a time.
client (OpenAI-compatible)
|
v
+---------------------------+ Harness Context Protocol (HCP)
| MoL Proxy |<--- versioned TOML contract: model and provider
| route -> answer -> summary| selection, tool policy, skills, prompts,
+---------------------------+ MCP servers, hooks, session state
|
v
+---------------------------+
| engine (vLLM / SGLang) | native multi-LoRA, adapters published as L0-L3
| |
| +-------------------+ |
| | frozen base | | 744B GLM-5.2 (Venti) | 35B-A3B Qwen3.6 (Tall)
| +-------------------+ |
| ^ ^ ^ ^ |
| L0 L1 L2 L3 | Chat | Agent | Coding | GenUI
+---------------------------+
|
v
tools, terminal, UI4A renderer (TSX: React / SolidJS)
Base. Frozen. No gradients ever reach it. Sharing one resident base is what makes the adapter population cheap to store and what lets a new specialist be registered without touching the weights every other specialist depends on.
LoRAs. Four release-labeled specialists resident in the engine under exactly
the names L0 through L3. L0 is Chat: conversational backbone, instruction
following, model identity. L1 is Agent: long-horizon tool use, personal-agent
workflows, service integrations. L2 is Coding: code generation, SWE-style tasks,
terminal use. L3 is GenUI: UI4A rendering and UI-driven action, specialized on
TSX with framework-specific renderers for other targets.
L0 Router. There is no separately trained router model. L0 classifies each
incoming user turn into exactly one of L0-L3 under a 24-token decode budget,
with constrained decoding restricting output to those four labels. The router
prompt treats the request as quoted untrusted text and checks wrapper families
in priority order: generative UI, then code and terminal, then personal-agent.
If L0 selects itself the turn stays on L0; otherwise the proxy switches.
Harness. Each fresh user request runs three hops: route, then answer on the selected adapter using its own view of the conversation, then a short cross-adapter summary retained proxy-side and never returned to the client. Tool results skip routing entirely and stay on whichever adapter issued the call. Per-adapter prompts are stable, so the engine's prefix cache reuses existing KV prefixes across turns. The runtime configuration is an HCP artifact: a declarative document carrying no gradients, but addressable, so a model version can propose edits to prompts, skills, tool allowlists, and hooks, have them evaluated on a re-run of the affected slice, and ship the survivors as the next configuration. The trainable object is the model-harness pair.
Tools and UI. UI4A is a component-native generative-UI harness; L3 targets it directly. Agent runtimes, plugin manifests, the local WebUI, and GenUI tooling live in macaron-artifacts. The HCP SDK is at hcp-sdk.
Reported as of the technical report, August 2026. Higher is better, all values
on a 0-100 scale. * marks a value imported from a public leaderboard or model
report rather than rerun by us.
| Benchmark | Venti | GLM-5.2 | GPT-5.5 | Opus 4.8 | Gemini 3.1 | Qwen 3.7 | Minimax M3 |
|---|---|---|---|---|---|---|---|
| Personal Intelligence | |||||||
| ChatBench | 58.3 | 54.5 | 55.5 | 52.8 | 52.0 | 52.5 | 49.1 |
| LivingBench | 64.0 | 60.5 | 61.9 | 63.8 | 52.1 | 56.1 | 57.1 |
| Agent | |||||||
| VitaBench | 60.0 | 55.8 | 55.8 | 56.5 | 55.2 | 61.2 | 56.8 |
| VitaBench2 | 46.0 | 43.1 | 47.4 | 46.3 | 50.2 | 47.6 | 39.4 |
| tau^3-Bench | 69.3 | 69.1 | 61.1 | 67.7 | 67.1* | 63.0 | 61.2 |
| PinchBench | 94.0 | 88.1 | 89.0* | 91.8* | 82.9* | 93.4* | 86.1 |
| ClawGym | 77.7 | 74.6 | 82.5 | 80.5 | 77.5 | 75.7 | 76.2 |
| Coding and terminal | |||||||
| SWE-Verified | 85.6 | 80.4 | 82.9* | 88.6* | 80.6* | 80.4* | 80.5* |
| TerminalBench 2.1 | 87.6 | 82.7* | 83.4* | 78.9* | 70.7* | 73.5* | 66.0* |
| DeepSWE | 58.4 | 54.9* | 70.0* | 58.0* | 10.0* | 18.0* | 20.0* |
| SWE Atlas QnA | 49.5 | 48.9* | 45.4* | 57.3* | 13.5* | 22.6 | 37.9 |
| GenUI | |||||||
| UI4A-Bench | 87.8 | 67.1 | 72.1 | 75.9 | 60.3 | 62.5 | 63.0 |
Tall against its own base, on the seven benchmarks evaluated for both systems under the same protocol within each row:
| Benchmark | Tall | Qwen3.6 35B-A3B |
|---|---|---|
| ChatBench | 54.9 | 48.0 |
| LivingBench | 48.4 | 47.1 |
| PinchBench | 86.2 | 82.5 |
| ClawGym | 64.0 | 58.6 |
| SWE-Verified | 75.4 | 73.4 |
| TerminalBench 2.1 | 56.2 | 52.5 |
| UI4A-Bench | 59.3 | 33.9 |
Within a benchmark row, every unstarred value uses the same task set and the same benchmark-specific protocol, which is what supports comparison on that row. Across rows, metrics, scaffolds, retry policies, and estimands differ, so the table does not define an aggregate ranking and we do not compute one.
Specifics that change how a row should be read: all VitaBench values are reruns under the same reproduced GLM-5.1 judge and user protocol. VitaBench2 uses Avg@1 under the Rewrite/Agentic Memory setting for every model, while the official leaderboard reports Avg@4. tau^3-Bench is pass@1. PinchBench for Venti is a best-observed score. SWE and terminal evaluations use the Claude Code agent scaffold rather than the production MoL harness. UI4A-Bench runs 161 cases in the shared UI4A runtime with adapter-free UI4A for baselines, and the primary score is the mobile viewport. Routing diagnostics come from a 6,448-sample trace drawn from LoRA training data, which is not an independent held-out split.
- No confidence intervals and no human-judge sensitivity study for the Personal Intelligence benchmarks. ChatBench and LivingBench source domains overlap product iteration and the self-improvement loop, so they measure the distribution this release targets rather than an independent one.
- Starred cells are not matched reruns and are excluded from any protocol-equivalent claim.
- The routing measurement is an implementation diagnostic. It does not estimate routing generalization.
- Diagnostics establish that an exercised path works. They do not establish component causality, continual improvement across generations, or collective intelligence across independently trained adapters. Whether composing adapters trained by different teams yields capability beyond any constituent specialist is unresolved in this release.
- Internal evaluations include de-identified product conversations and traffic. The report does not document the consent basis for research use, the de-identification procedure, a residual re-identification audit, retention and access controls, or an ethics-review determination.
- There is no standalone safety and red-team evaluation, and no complete per-specialist training specification. Treat these results as a systems characterization, not as evidence of suitability for safety-critical use.
Reproduction artifacts are not public. The harness repository ships serving code only; weights, benchmark cases, and test suites are not included in it, and no evaluation scripts are released alongside this snapshot.
OpenAI-compatible. https://mintcn.macaron.xin/ for China Mainland,
https://mint.macaron.im/ for the internal trial. Self-hosted, the MoL proxy
exposes GET /health, GET /v1/models, POST /v1/chat/completions, and
POST /v1/responses. Chat Completions is stateless with client-resent history;
Responses keeps state proxy-side and resumes from previous_response_id.
Validated operating points from the report, rather than minimum requirements:
| Hardware | Engine | Parallelism | Attention | Context | Role |
|---|---|---|---|---|---|
| H20 | vLLM 0.24 | TP4/PP2/DCP4 | FLASHMLA_SPARSE / fp8_ds_mla | 262K | Venti reference |
| B300 | vLLM 0.24 | TP8/DCP4 | FLASHMLA_SPARSE / auto | 1M | Single-node fallback |
| B300 | SGLang 0.5.15.post1 | PCP + CP LayerSplit prefill, DCP4 + EAGLE decode | FlashMLA sparse, L3 HiCache | 1M | Production RDMA PD worker |
Neither model card publishes a VRAM figure. The checked-in harness profiles
assume eight visible GPUs and share max_num_seqs=8,
gpu_memory_utilization=0.915, prefix caching, and CUDA graphs to 8. Observed
concurrency on H20 with TP4/PP2/DCP4: sixteen 56K-token requests, eight
180K-token requests, or four 230K-token requests. B300 DCP2/DCP4/DCP8 with EAGLE
provide approximately 2.34M, 4.67M, and 9.34M logical KV tokens. Serving
efficiency depends on the joint choice of engine, attention backend, parallelism
layout, speculative decoding, and load; treat each row as one scoped operating
point rather than a portable configuration.
No quantized weights are released. Both models ship BF16 base checkpoints, and
Tall's L2 adapter is stored in F32. FP8 appears in the validated profiles as a
serving-time choice, fp8_ds_mla attention on H20 and FP8 KV cache at page size
64 on B300, not as a released weight format. Community quantizations exist on
Hugging Face and are not produced or verified by us.
The MoL proxy has no access control of its own. Production paths expect an API
key through MOL_API_KEY_FILE. The MOL_ALLOW_UNAUTHENTICATED=1 escape hatch
is intended for an isolated development host; a proxy started that way and
reachable from a network is an open inference endpoint. The L0 router prompt
frames incoming requests as quoted untrusted text, but that is a routing
safeguard, not a general defense against prompt injection reaching the tool
surface. Scope tool allowlists in the HCP artifact accordingly.
Model weights and the Mixture-of-LoRA-Harness are MIT. Venti additionally
inherits the terms of its GLM-5.2 base; Tall inherits the terms of its
Qwen3.6-35B-A3B base. macaron-artifacts is Apache-2.0. Harness dependencies
carry their own terms.
@misc{mindlab2026macaronv1,
title = {Macaron-V1: Towards Open Continual Learning with Self-Improvement
and Mixture-of-LoRA},
author = {{Mind Lab}},
year = {2026},
eprint = {2608.09819},
url = {https://huggingface.co/papers/2608.09819}
}