Link
https://github.com/Fr0zenCrane/jev-spatial
Live site (optional)
No response
What does it do, and where does Jev fit?
Jev-Spatial is a System One spatial model inspired by Jev. It builds on Molmo2-ER, and instead of generating text it answers spatial questions by choosing from a fixed set of options with one shared classifier head. Classification tasks such as spatial relations, directions and yes/no take one choice. Numeric regression takes two: a value range, then a sub-range. Pointing is a special case: three rounds of choosing a cell in a 3×3 grid, cropping into the chosen cell each time. Because the model only picks options, every output is valid and nothing needs parsing. Overall it performs about as well as its base model: within 1–4 points on classification, with lower error on metric estimation. Pointing was a nice surprise and actually beats the base model (+2.0 on RefSpatial, +7.0 on Where2Place). Questions about the same image share one image encoding and run in parallel, so with 8 questions per image it is about 4.2× faster than answering them one at a time with an autoregressive model. Code and weights are open, and feedback is welcome!
Category
Model & Train / Infer Code
Link
https://github.com/Fr0zenCrane/jev-spatial
Live site (optional)
No response
What does it do, and where does Jev fit?
Jev-Spatial is a System One spatial model inspired by Jev. It builds on Molmo2-ER, and instead of generating text it answers spatial questions by choosing from a fixed set of options with one shared classifier head. Classification tasks such as spatial relations, directions and yes/no take one choice. Numeric regression takes two: a value range, then a sub-range. Pointing is a special case: three rounds of choosing a cell in a 3×3 grid, cropping into the chosen cell each time. Because the model only picks options, every output is valid and nothing needs parsing. Overall it performs about as well as its base model: within 1–4 points on classification, with lower error on metric estimation. Pointing was a nice surprise and actually beats the base model (+2.0 on RefSpatial, +7.0 on Where2Place). Questions about the same image share one image encoding and run in parallel, so with 8 questions per image it is about 4.2× faster than answering them one at a time with an autoregressive model. Code and weights are open, and feedback is welcome!
Category
Model & Train / Infer Code