Multimodal Perception
Perception
Seeing the world as one signal
Vision, audio and language fused into a single latent space so agents understand scenes the way people do.
An independent AI lab shaping foundation models, perception systems and the interfaces of tomorrow.
Where curiosity is a discipline.
We run long-horizon research programs across perception, reasoning and alignment — then ship what survives contact with reality.

Seeing the world as one signal
Vision, audio and language fused into a single latent space so agents understand scenes the way people do.

From answers to actions
Long-horizon planning, tool use and self-verification for agents that operate reliably over hours, not seconds.

Capability without surprises
Interpretability, red-teaming and scalable oversight baked into every training run.

Frontier quality, edge budget
Distillation, quantisation and custom kernels that bring frontier models to phones, cars and factories.
A family of foundation models, tuned for production.
Every model ships with evaluation reports, safety cards and a latency budget. Deploy in our cloud, your VPC, or on the edge.
Frontier multimodal foundation model
Our largest model. Native Arabic and English reasoning, tool use and agentic workflows with a one-million-token memory.
Real-time perception for the physical world
Detection, tracking, segmentation and document understanding at video frame-rate on a single GPU.
Frontier reasoning, on-device
A 3B-parameter model distilled from Core that runs offline on phones and embedded hardware with full Arabic support.
Systems in the wild.
A few of the deployments we are allowed to talk about.
Tell us what you are working on.
Partnerships, research collaborations, model access or just a conversation — our inbox is open.