Multiverse Computing Opens Its Compressed AI Models to Developers
Spanish AI startup Multiverse Computing launched a self-serve API portal and a consumer chat app built on its quantum-inspired compressed models, targeting enterprises that need production-ready AI without AWS Marketplace friction. The move puts compressed model inference—faster, cheaper, deployable on-device—directly in developers' hands.
Multiverse Computing, the Spanish startup that compresses models from OpenAI, Meta, DeepSeek, and Mistral AI into leaner, faster versions, has stopped requiring an AWS Marketplace account to get access to its technology. As of March 19, developers and enterprises can hit a self-serve API portal directly—real-time usage monitoring included, no procurement workflow required.
The same day, the company launched CompactifAI, a consumer-facing chat app that demos its compressed models against the originals. It’s ChatGPT-shaped on the surface, but the actual pitch is infrastructure: Multiverse wants businesses to see that a compressed Llama or Mistral variant can handle their workload for materially less compute spend.
Compression via quantum-inspired tensor network methods is the core IP. The technique reduces parameter count and memory footprint while preserving accuracy on domain-specific tasks—often better than the base model on narrow benchmarks because the compression is applied after fine-tuning on client data. Multiverse’s HyperNova 60B, released for free on Hugging Face in February, demonstrated that a compressed 60B model can outperform some uncompressed 70B models on reasoning tasks.
The timing is deliberate. Enterprise AI adoption is stalling at proof-of-concept because running large frontier models at scale costs more than most IT budgets expected. A compressed model that delivers 85% of GPT-5.4’s performance at 30% of the cost is a real conversation at the CFO level, and Multiverse is positioning its API portal as where that conversation becomes a contract.
The strategic backdrop: Multiverse is reportedly seeking a €500M funding round. Opening a direct API channel—with usage data flowing back—gives the company both revenue and the traction metrics that justify that valuation. The Axelera AI partnership announced March 18, which targets edge hardware deployment, extends the same argument to on-device inference, where model size is an absolute constraint.
Compression remains a second-class citizen in most enterprise AI discussions, dominated by the compute-scaling narratives of OpenAI and Google. But for verticals where data residency, latency, and recurring inference costs matter—financial services, healthcare, defense—a smaller, faster model that runs in your own infrastructure is often the right answer. Multiverse is betting the enterprise market figures that out in 2026.