A new model class
Featherweights
The AI frontier won’t be won with the biggest resources but with the smartest models.
Today
The industry is
chasing the ratio.
Frontier AI is bottlenecked on compute: roughly $725B of infrastructure this year, gigawatt campuses included, and still not enough. That makes any lever on capability per unit of compute steadily more valuable. The clearest signal yet — NVIDIA, the company selling the compute, just put a reported $5B into Safe Superintelligence, partly for access to research it expects will get more out of its chips.
NVIDIA · Reuters · Bloomberg · Epoch AI · July 2026.
Epoch AI · compute of notable AI models
Tomorrow
The biggest lever
is the model class.
Every frontier model is a transformer, and every efficiency gain in production — quantization, distillation, sparsity, faster kernels — is a tune-up of the same architecture. Each moves the ratio at the margin. None changes what the model class fundamentally costs.
That is the lever we pulled. Ours is an architecture developed before the transformer’s success — a different foundation, not another tune-up.
At small scale the advantage is already measurable: more capability per unit of compute, at far lower cost to engineer, train, and run. And something more: even at this scale, a completely new form of intelligence appears to be emerging from it.
Our solution
A new primitive
for intelligence.
Today’s frontier models are built from weights: billions of floating-point numbers, multiplied at massive scale and held together by a growing stack of hand-engineered techniques.
We believe we have found a fundamentally different building block: a new computing primitive, trained and run in a very different and far more efficient manner than current models — even those at the frontier of research.
The numbers
The results so far are extremely promising.
Each model's officially published benchmarks beside ours, and the hardware each needs to run. Comparable capability, far less compute.
Benchmarks higher is better
SWE-bench Verifiedagentic coding
LiveCodeBench v6code gen
GPQA Diamondreasoning
BrowseCompweb agent
HLEfrontier exam
GLM-4.7-Flash30B MoE, 1-bit
- GPU
- L4 24GB
- vCPU
- 8
- RAM
- 32 GB
MiniMax-M2.7229B MoE, 1-bit
- GPU
- A100 80GB
- vCPU
- 24
- RAM
- 220 GB
TysonMiniMax-M2.7 distilled
- GPU
- TBD
- vCPU
- TBD
- RAM
- TBD
Resource advantage over comparable models projected with scale
Rocky already runs on a small fraction of the resources a comparable model needs. We expect an efficiency advantage of around 100x or more at the frontier, as that lead only widens with scale.
Methodology: We compare against 1-bit quantized models on purpose, to show these advantages are independent of known efficiency techniques like quantization: we reach comparable capability at a fraction of the compute, rather than trading capability away for size. Rocky and Tyson are also distillations, which removes training data as a confound. Distillation under these constraints would ordinarily be expected to reduce capability sharply, yet they stay highly capable, which we read as evidence that the advantage lies in the architecture itself.
The round
We are raising a round of up to
$5M
Minimum contribution · $100k
This round funds the decisive experiment: scaling the architecture against well-tuned baselines at matched compute. Its job is to prove the model at much larger scales — and that proof sets up the next stage: a much, much larger raise.
Get in touch
We are moving
with conviction.
If this is something you want to be part of, in whatever capacity, we would welcome a conversation, and the earlier the better.
contact@featherweights.ai · This page is private — please do not circulate.