Skip to content

Future: compact decision model (<10M params) #59

Description

@Oaklight

Parent: #58

Idea: Purpose-built tiny decision model — 2-4 transformer layers, 128-256 width, byte-level or subword encoder.

References:

  • CUA-S1: 706K params, 2-layer transformer, byte encoder → 99.7% on form-filling
  • jevlike: AttentionHead with byte encoder
  • Ettin-68M: 68M params, 28ms — current smallest viable backbone

Open questions:

  • Can a <10M model match Ettin-68M's zero-shot 47.6% avg accuracy?
  • Does a byte encoder avoid tokenizer-dependent surface-form bias?
  • What's the minimum capacity for each head type (noul is simpler than choice)?

Why deferred: Ettin-68M at 28ms is already fast enough for most use cases. Building a custom architecture from scratch is high effort for marginal latency gains.

When to revisit: If there's demand for <10ms inference or edge deployment.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:modelDecision model replication (src, training, evaluation)area:researchSurvey, literature review, ecosystem analysisdeferredFuture exploration, not on current roadmap

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions