Geometry Foundation Model.
One model that grounds perception in 3D: reconstructs from real captures, holds geometry, viewpoint, scale and occlusion, and stays consistent when the camera moves. The substrate everything else stands on.

Computer vision still has no foundation model — no geometry foundation model, no visual general intelligence. We are building both, in the open.
Embedding-guided Neural Ensemble for Adaptive Segmentation
Name a concept in plain language — ENEAS finds it, segments it, and holds onto that exact instance across an entire image or video.
Read the researchComputer vision — how a machine understands the real world, its space and its 3D structure — and 3D generation are treated as two fields. They are one problem. Understanding a world and generating a world depend on the same missing layer: a coherent internal representation of space. A model that cannot hold space cannot reliably understand it, and cannot reliably generate it either. Language already went through this: the model that predicts text turned out to be the model that understands it.
Vision has no such foundation yet — no Geometry Foundation Model (GFM) that grounds perception in 3D, and no Visual General Intelligence (VGI) that reasons and generates over it. 3D understanding and 3D generation are not solved. That is the research programme, and we run it in the open.
One model that grounds perception in 3D: reconstructs from real captures, holds geometry, viewpoint, scale and occlusion, and stays consistent when the camera moves. The substrate everything else stands on.
Reasoning, editing, simulation and generation as different queries to the same internal world — perception and generation in one model, the way a single language model both reads and writes.
Weights, evaluations, and methods, released as we go. We would rather move the field forward in the open than keep results behind a wall — so anyone can build on what we ship.