Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 

Repository files navigation

Macaron-V1

An open agent-model family built for experiential intelligence: learning from experience accumulated in a real environment, and continuing to learn after deployment.

Try API | Run locally | HF Models | Technical report | Evaluation


Macaron-V1 ships two releases that share one architecture. Both keep a large base frozen, layer four LoRA specialists on top of it (Chat, Agent, Coding, GenUI), and select exactly one specialist per user turn through a shared runtime. The specialists are portable: new ones can be registered against the same base without retraining it, which is the substrate for continual learning and for composing adapters trained by different teams.

Venti is the flagship. Tall is the same four-specialist design on a smaller base for local and lower-latency deployment. Neither is a fine-tune of the other; they are separate bases running the same Mixture-of-LoRA (MoL) design.

Choose a model

Macaron-V1-Venti Macaron-V1-Tall
Release label 748B 50B
Composition 744B frozen GLM-5.2 base + 4 LoRA specialists 35B-A3B frozen Qwen3.6 base + 4 LoRA specialists
Adapter config rank 16, alpha 32, 7,688,042,496 stored values each rank 64, alpha 128, 3,775,651,840 stored values each (L2 stored in F32)
Context 1M on B300-class configurations; 262K in the reference vLLM profile 262K
Best fit Frontier-quality personal agent, long-horizon tool use, GenUI, SWE and terminal work Local or on-prem deployment, latency-sensitive serving, lower cost per request
Deployment 8 GPU (H20 or B300 class) via vLLM or SGLang, plus the MoL proxy 8 GPU reference profile via vLLM, plus the MoL proxy
Cost profile One resident base serves all four specialists; ~774.8B logical parameters resident against 2.976T for four separately merged base copies Same sharing property at roughly 1/15 the base footprint
Weights mindlab-research/Macaron-V1-Venti mindlab-research/Macaron-V1-Tall

A third build, Macaron-V1-Coding-Venti, merges the coding LoRA directly into the base instead of serving it as an adapter. It is excluded from every evaluation claim below.

Release labels are release-facing, not residency counts. Venti's four public adapter headers sum to about 30.8B stored values, so the stored total is approximately 774.8B rather than 748B. Tall's Hugging Face metadata reports 36B because it counts the base checkpoint alone; the 50B label counts base plus adapters, which is the accounting basis that makes it comparable to Venti.

Highlights

We are releasing the model weights and the technical report for Macaron-V1, a 748B open-source agentic model built for personal intelligence and continual learning. Alongside the weights we are releasing the methodology and infrastructure behind it: the AutoResearch platform, Model-Harness Co-design, and the Macaron ChatBench and LivingBench suites.

Three key points:

  • Releasing the model weights and technical report of Macaron V1.
  • Macaron V1 is our deliberate experiment: a 748B open-source agentic model designed for personal intelligence and continual learning.
  • A novel model architecture, Mixture of LoRA, routes requests across four domain experts over a frozen base, not just fine-tuning.
  • Alongside Macaron V1, we are releasing the methodology and infrastructure behind it: the infrastructure and platform of AutoResearch; Model-Harness Co-design; Macaron ChatBench and Macaron LivingBench.

Quickstart

The hosted API is OpenAI-compatible.

Endpoint Base URL
China Mainland https://mintcn.macaron.xin/
Internal trial https://mint.macaron.im/

To serve the weights yourself, use the Mixture-of-LoRA-Harness, which supplies the L0 router prompt, per-adapter conversation views, and same-request adapter switching in front of a vLLM or SGLang engine. The harness requires Linux with CUDA, Python 3.10+, eight visible GPUs for the checked-in profiles, and either vLLM 0.24.0 or SGLang 0.5.13.post1 installed separately against your host CUDA stack. Weights are not bundled with the harness; pull them from Hugging Face and place the base shards and loras/L0-L3 under the model root the launch scripts expect.

Validate a profile before starting an engine:

VLLM_CONFIG_CHECK_ONLY=1 ./mol_harness/scripts/start_vllm_venti.sh

Then bring up the engine and proxy together. The proxy listens on 8200 and the engine on 127.0.0.1:8000:

MOL_API_KEY_FILE=/path/to/api-key ./mol_harness/scripts/restart_vllm_tp8_venti.sh

Swap venti for tall to run the smaller profile. Both profiles target the same GPUs and ports, so only one runs at a time.

Architecture

  client (OpenAI-compatible)
          |
          v
  +---------------------------+     Harness Context Protocol (HCP)
  |        MoL Proxy          |<--- versioned TOML contract: model and provider
  |  route -> answer -> summary|     selection, tool policy, skills, prompts,
  +---------------------------+     MCP servers, hooks, session state
          |
          v
  +---------------------------+
  |   engine (vLLM / SGLang)  |  native multi-LoRA, adapters published as L0-L3
  |                           |
  |   +-------------------+   |
  |   |   frozen base     |   |  744B GLM-5.2 (Venti) | 35B-A3B Qwen3.6 (Tall)
  |   +-------------------+   |
  |     ^     ^     ^     ^   |
  |    L0    L1    L2    L3   |  Chat | Agent | Coding | GenUI
  +---------------------------+
                |
                v
      tools, terminal, UI4A renderer (TSX: React / SolidJS)

Base. Frozen. No gradients ever reach it. Sharing one resident base is what makes the adapter population cheap to store and what lets a new specialist be registered without touching the weights every other specialist depends on.

LoRAs. Four release-labeled specialists resident in the engine under exactly the names L0 through L3. L0 is Chat: conversational backbone, instruction following, model identity. L1 is Agent: long-horizon tool use, personal-agent workflows, service integrations. L2 is Coding: code generation, SWE-style tasks, terminal use. L3 is GenUI: UI4A rendering and UI-driven action, specialized on TSX with framework-specific renderers for other targets.

L0 Router. There is no separately trained router model. L0 classifies each incoming user turn into exactly one of L0-L3 under a 24-token decode budget, with constrained decoding restricting output to those four labels. The router prompt treats the request as quoted untrusted text and checks wrapper families in priority order: generative UI, then code and terminal, then personal-agent. If L0 selects itself the turn stays on L0; otherwise the proxy switches.

Harness. Each fresh user request runs three hops: route, then answer on the selected adapter using its own view of the conversation, then a short cross-adapter summary retained proxy-side and never returned to the client. Tool results skip routing entirely and stay on whichever adapter issued the call. Per-adapter prompts are stable, so the engine's prefix cache reuses existing KV prefixes across turns. The runtime configuration is an HCP artifact: a declarative document carrying no gradients, but addressable, so a model version can propose edits to prompts, skills, tool allowlists, and hooks, have them evaluated on a re-run of the affected slice, and ship the survivors as the next configuration. The trainable object is the model-harness pair.

Tools and UI. UI4A is a component-native generative-UI harness; L3 targets it directly. Agent runtimes, plugin manifests, the local WebUI, and GenUI tooling live in macaron-artifacts. The HCP SDK is at hcp-sdk.

Evaluation

Reported as of the technical report, August 2026. Higher is better, all values on a 0-100 scale. * marks a value imported from a public leaderboard or model report rather than rerun by us.

Benchmark Venti GLM-5.2 GPT-5.5 Opus 4.8 Gemini 3.1 Qwen 3.7 Minimax M3
Personal Intelligence
ChatBench 58.3 54.5 55.5 52.8 52.0 52.5 49.1
LivingBench 64.0 60.5 61.9 63.8 52.1 56.1 57.1
Agent
VitaBench 60.0 55.8 55.8 56.5 55.2 61.2 56.8
VitaBench2 46.0 43.1 47.4 46.3 50.2 47.6 39.4
tau^3-Bench 69.3 69.1 61.1 67.7 67.1* 63.0 61.2
PinchBench 94.0 88.1 89.0* 91.8* 82.9* 93.4* 86.1
ClawGym 77.7 74.6 82.5 80.5 77.5 75.7 76.2
Coding and terminal
SWE-Verified 85.6 80.4 82.9* 88.6* 80.6* 80.4* 80.5*
TerminalBench 2.1 87.6 82.7* 83.4* 78.9* 70.7* 73.5* 66.0*
DeepSWE 58.4 54.9* 70.0* 58.0* 10.0* 18.0* 20.0*
SWE Atlas QnA 49.5 48.9* 45.4* 57.3* 13.5* 22.6 37.9
GenUI
UI4A-Bench 87.8 67.1 72.1 75.9 60.3 62.5 63.0

Tall against its own base, on the seven benchmarks evaluated for both systems under the same protocol within each row:

Benchmark Tall Qwen3.6 35B-A3B
ChatBench 54.9 48.0
LivingBench 48.4 47.1
PinchBench 86.2 82.5
ClawGym 64.0 58.6
SWE-Verified 75.4 73.4
TerminalBench 2.1 56.2 52.5
UI4A-Bench 59.3 33.9

Protocol

Within a benchmark row, every unstarred value uses the same task set and the same benchmark-specific protocol, which is what supports comparison on that row. Across rows, metrics, scaffolds, retry policies, and estimands differ, so the table does not define an aggregate ranking and we do not compute one.

Specifics that change how a row should be read: all VitaBench values are reruns under the same reproduced GLM-5.1 judge and user protocol. VitaBench2 uses Avg@1 under the Rewrite/Agentic Memory setting for every model, while the official leaderboard reports Avg@4. tau^3-Bench is pass@1. PinchBench for Venti is a best-observed score. SWE and terminal evaluations use the Claude Code agent scaffold rather than the production MoL harness. UI4A-Bench runs 161 cases in the shared UI4A runtime with adapter-free UI4A for baselines, and the primary score is the mobile viewport. Routing diagnostics come from a 6,448-sample trace drawn from LoRA training data, which is not an independent held-out split.

Limitations

  • No confidence intervals and no human-judge sensitivity study for the Personal Intelligence benchmarks. ChatBench and LivingBench source domains overlap product iteration and the self-improvement loop, so they measure the distribution this release targets rather than an independent one.
  • Starred cells are not matched reruns and are excluded from any protocol-equivalent claim.
  • The routing measurement is an implementation diagnostic. It does not estimate routing generalization.
  • Diagnostics establish that an exercised path works. They do not establish component causality, continual improvement across generations, or collective intelligence across independently trained adapters. Whether composing adapters trained by different teams yields capability beyond any constituent specialist is unresolved in this release.
  • Internal evaluations include de-identified product conversations and traffic. The report does not document the consent basis for research use, the de-identification procedure, a residual re-identification audit, retention and access controls, or an ethics-review determination.
  • There is no standalone safety and red-team evaluation, and no complete per-specialist training specification. Treat these results as a systems characterization, not as evidence of suitability for safety-critical use.

Reproduction artifacts are not public. The harness repository ships serving code only; weights, benchmark cases, and test suites are not included in it, and no evaluation scripts are released alongside this snapshot.

Deploy and trust

API

OpenAI-compatible. https://mintcn.macaron.xin/ for China Mainland, https://mint.macaron.im/ for the internal trial. Self-hosted, the MoL proxy exposes GET /health, GET /v1/models, POST /v1/chat/completions, and POST /v1/responses. Chat Completions is stateless with client-resent history; Responses keeps state proxy-side and resumes from previous_response_id.

Local and cluster

Validated operating points from the report, rather than minimum requirements:

Hardware Engine Parallelism Attention Context Role
H20 vLLM 0.24 TP4/PP2/DCP4 FLASHMLA_SPARSE / fp8_ds_mla 262K Venti reference
B300 vLLM 0.24 TP8/DCP4 FLASHMLA_SPARSE / auto 1M Single-node fallback
B300 SGLang 0.5.15.post1 PCP + CP LayerSplit prefill, DCP4 + EAGLE decode FlashMLA sparse, L3 HiCache 1M Production RDMA PD worker

Neither model card publishes a VRAM figure. The checked-in harness profiles assume eight visible GPUs and share max_num_seqs=8, gpu_memory_utilization=0.915, prefix caching, and CUDA graphs to 8. Observed concurrency on H20 with TP4/PP2/DCP4: sixteen 56K-token requests, eight 180K-token requests, or four 230K-token requests. B300 DCP2/DCP4/DCP8 with EAGLE provide approximately 2.34M, 4.67M, and 9.34M logical KV tokens. Serving efficiency depends on the joint choice of engine, attention backend, parallelism layout, speculative decoding, and load; treat each row as one scoped operating point rather than a portable configuration.

Quantization

No quantized weights are released. Both models ship BF16 base checkpoints, and Tall's L2 adapter is stored in F32. FP8 appears in the validated profiles as a serving-time choice, fp8_ds_mla attention on H20 and FP8 KV cache at page size 64 on B300, not as a released weight format. Community quantizations exist on Hugging Face and are not produced or verified by us.

Security

The MoL proxy has no access control of its own. Production paths expect an API key through MOL_API_KEY_FILE. The MOL_ALLOW_UNAUTHENTICATED=1 escape hatch is intended for an isolated development host; a proxy started that way and reachable from a network is an open inference endpoint. The L0 router prompt frames incoming requests as quoted untrusted text, but that is a routing safeguard, not a general defense against prompt injection reaching the tool surface. Scope tool allowlists in the HCP artifact accordingly.

License

Model weights and the Mixture-of-LoRA-Harness are MIT. Venti additionally inherits the terms of its GLM-5.2 base; Tall inherits the terms of its Qwen3.6-35B-A3B base. macaron-artifacts is Apache-2.0. Harness dependencies carry their own terms.

Citation

@misc{mindlab2026macaronv1,
  title  = {Macaron-V1: Towards Open Continual Learning with Self-Improvement
            and Mixture-of-LoRA},
  author = {{Mind Lab}},
  year   = {2026},
  eprint = {2608.09819},
  url    = {https://huggingface.co/papers/2608.09819}
}

Contact

contact@mindlab.ltd

About

An open agent-model family built for experiential intelligence

Topics

Resources

Stars

5 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors