Beyond Domain-Specific World Models: JEPA-Anything Uses 1 Recipe for 7 Fields


Researchers from PhAI Labs, CUHK, Fudan, Stanford, Oxford and Princeton have released JEPA-Anything, a domain-agnostic framework for building world models. Instead of designing a new predictive model for each field, it applies one shared learning recipe to very different systems. It extends joint-embedding predictive architectures (JEPAs) with a method called Orthogonal Predictive Factorization (OPF). The research team tested it across 7 domains: vision, biology, clinical trajectories, control, molecular dynamics, physical fields and weather.

What problem does JEPA-Anything solve?

A standard JEPA, such as I-JEPA or V-JEPA 2, uses a context encoder, an EMA target encoder and one predictor. The predictor outputs one monolithic target embedding. The research team call this a capacity-allocation problem: high-variance structure dominates, and weaker modes get conflicting gradients.

How does Orthogonal Predictive Factorization work?

OPF splits the latent target of width d into K learned subspaces of width r, with d = K × r. Most experiments use K = 4. Each factor gets a dedicated predictor. The factor predictions are then recombined through the Moore-Penrose pseudoinverse of the projector matrix. The result is 1 complete latent state for decoding, planning or rollout.

Three regularizers keep the factors useful:

  • Orthogonality loss: keeps columns within each projector orthonormal and different projectors in non-overlapping subspaces.
  • Factor-activity loss: a hinge on per-coordinate standard deviation so no factor goes dead.
  • Encoder-variance loss: sends a direct anti-collapse signal to the online encoder.

The OPF loss is simply added to each domain’s original training loss. Domain adapters handle tokenization and encoders; the core library exposes the shared core as OrthogonalFactorProjection.

Orthogonality matters for stable synthesis. On CITRIS Interventional Pong, a capacity-matched unconstrained multi-head model had a condition number of 438.52. The orthogonal version reached 1.00005, with cross-factor overlap near zero.



Source link

  • Related Posts

    Reflection AI Introduces Beam: A 501B Open-Weight MoE Model With 23B Active Parameters for Coding and Agentic Workloads

    Reflection AI has introduced Beam, its first open-weight model. Beam is a sparse Mixture-of-Experts (MoE) model with 501B total parameters and 23B active per token, built for coding, reasoning and…

    Meet Together Link: A Free CLI That Runs Open Models Like Kimi K3 and GLM 5.3 Inside Claude Code, Codex, and OpenCode

    Together AI has released Together Link, a free, MIT-licensed CLI now in beta. It connects the coding agents developers already use to open models hosted on Together AI. Supported tools…

    Leave a Reply

    Your email address will not be published. Required fields are marked *