
Infinity raises $15M for CUDA-alternative inference software that works across any AI chip
Published by AINave Editorial • Reviewed by Ramit
Infinity, an AI infrastructure startup, has raised $15 million at a $100 million valuation to build a CUDA-alternative kernel stack that lets any AI chip run inference workloads efficiently. The company's Ignition AI research agent automates the writing, testing, and optimization of low-level code, potentially giving AI builders more hardware choices beyond Nvidia.
What happened
Infinity announced a $15 million seed round led by Touring Capital with participation from Principal VC and researchers from OpenAI and Anthropic. The company is developing a CUDA-alternative kernel stack that works across GPUs, SRAM, phone chips, and Systolic Arrays. Its Ignition AI research agent writes, tests, debugs, and automatically rewrites low-level code for AI inference on non-Nvidia chips. The system is self-optimizing and adapts to different chip architectures.
Founded last year by Jeremy Nixon, a former Google Brain researcher and creator of AGI House, Infinity aims to build a universal inference library that replicates state-of-the-art results across chip architectures. The company does not charge upfront license fees. Instead, it takes a share of performance gains and cost savings measured in tokens-per-second. Early customers include AI chip maker D-Matrix, and Infinity is in talks with other chip and cloud companies. The team has 26 employees.
Why AI builders should care
If Infinity delivers on its claims, AI builders could gain a practical path to run inference on hardware other than Nvidia GPUs without rewriting kernels from scratch. The cross-chip inference tooling could reduce lock-in to Nvidia's CUDA ecosystem, which currently dominates because PyTorch and TensorFlow are built on top of it. A universal inference library would let developers deploy models on alternative chips with less engineering overhead.
The Ignition agent's ability to automate low-level optimization could also accelerate hardware-software co-design. Instead of months of manual kernel tuning, the agent claims to reduce the process to hours or days. For teams building AI products that need to optimize inference cost or latency, this could open up new hardware options and pricing models.
The involvement of researchers from OpenAI and Anthropic as investors signals that leading AI labs see value in hardware-agnostic inference tooling. It also suggests potential alignment with safety and capability research in future hardware stacks.
Practical implications
Infinity's pay-for-performance model is a notable shift. Instead of paying upfront license fees, customers share a portion of the performance gains and cost savings, measured in tokens-per-second. This aligns Infinity's incentives with real-world inference efficiency improvements. If the model works, it could change how AI teams evaluate inference infrastructure costs.
For chipmakers like D-Matrix, Infinity's software could lower the barrier to entry for running popular AI models. That could increase competition in the AI chip market, potentially driving down inference costs for builders. The presence of OpenAI and Anthropic researchers among backers adds credibility and may attract more customers.
Caveats
Infinity's claims about CUDA-level performance and universal chip compatibility have not been independently verified in publicly available sources. The company is early-stage with only 26 employees and one named customer. The token-per-second cost-sharing model requires real-world validation across multiple chip environments. Product capabilities, pricing, and timelines may change as the company develops. AI builders should treat the announced capabilities as aspirational until independent benchmarks emerge.
FAQs
Sources
- Inference startup Infinity raises $15M from Touring Capital, OpenAI and Athropic researchers
- OpenAI | Research & Deployment
- One year at Anthropic, then $200M at $1B: The researchers who just...
- OpenAI.fm
- Researchers launch new AI startup, part of growing neolab... | LinkedIn
- OpenAI, Google AI researchers back Anthropic's Pentagon lawsuit
- Inference startup Infinity raises $15M from Touring Capital, OpenAI and Athropic researchers
- Infinity Raises $15 Million in Seed Funding to Build the Software Layer That Makes Any AI Chip Inference-Ready
- Claude may help someone build bombs and bioweopons, so Athropic...
- OpenAI and Athropic will kill apps way faster than I thought!
- Infinity Raises $15 Million in Seed Funding to Build the Software Layer That Makes Any AI Chip Inference-Ready
- Inference startup Infinity raises $15M fr... - aVenture News
- Infinity raises $15M for inference from Touring Capital ...
- Infinity Raises $15 Million in Seed Funding to Build the ...
- Infinity raises $15M from Touring Capital, OpenAI and ...






















