featherweights.ai

A new model class

Featherweights

The AI frontier won’t be won with the biggest resources but with the smartest models.

Today

The industry is
chasing the ratio.

Frontier AI is bottlenecked on compute: roughly $725B of infrastructure this year, gigawatt campuses included, and still not enough. That makes any lever on capability per unit of compute steadily more valuable. The clearest signal yet — NVIDIA, the company selling the compute, just put a reported $5B into Safe Superintelligence, partly for access to research it expects will get more out of its chips.

NVIDIA · Reuters · Bloomberg · Epoch AI · July 2026.

Tomorrow

The biggest lever
is the model class.

Every frontier model is a transformer, and every efficiency gain in production — quantization, distillation, sparsity, faster kernels — is a tune-up of the same architecture. Each moves the ratio at the margin. None changes what the model class fundamentally costs.

That is the lever we pulled. Ours is an architecture developed before the transformer’s success — a different foundation, not another tune-up.

At small scale the advantage is already measurable: more capability per unit of compute, at far lower cost to engineer, train, and run. And something more: even at this scale, a completely new form of intelligence appears to be emerging from it.

Our solution

A new primitive
for intelligence.

Today’s frontier models are built from weights: billions of floating-point numbers, multiplied at massive scale and held together by a growing stack of hand-engineered techniques.

We believe we have found a fundamentally different building block: a new computing primitive, trained and run in a very different and far more efficient manner than current models — even those at the frontier of research.

The numbers

The results so far are extremely promising.

Each model's officially published benchmarks beside ours, and the hardware each needs to run. Comparable capability, far less compute.

Benchmarks higher is better

GLM-4.7-Flash Rocky MiniMax-M2.7
59.256n/r

SWE-bench Verifiedagentic coding

6461n/r

LiveCodeBench v6code gen

75.28689.8

GPQA Diamondreasoning

42.87477.8

BrowseCompweb agent

14.42628

HLEfrontier exam

GLM-4.7-Flash30B MoE, 1-bit

GPU
L4 24GB
vCPU
8
RAM
32 GB

RockyGLM-4.7-Flash distilled

GPU
None
vCPU
2 (shared)
RAM
624 MB
Request access

MiniMax-M2.7229B MoE, 1-bit

GPU
A100 80GB
vCPU
24
RAM
220 GB
In training

TysonMiniMax-M2.7 distilled

GPU
TBD
vCPU
TBD
RAM
TBD
100x+ today at scale

Resource advantage over comparable models projected with scale

Rocky already runs on a small fraction of the resources a comparable model needs. We expect an efficiency advantage of around 100x or more at the frontier, as that lead only widens with scale.

Methodology: We compare against 1-bit quantized models on purpose, to show these advantages are independent of known efficiency techniques like quantization: we reach comparable capability at a fraction of the compute, rather than trading capability away for size. Rocky and Tyson are also distillations, which removes training data as a confound. Distillation under these constraints would ordinarily be expected to reduce capability sharply, yet they stay highly capable, which we read as evidence that the advantage lies in the architecture itself.

The round

We are raising a round of up to

$5M

Minimum contribution · $100k

This round funds the decisive experiment: scaling the architecture against well-tuned baselines at matched compute. Its job is to prove the model at much larger scales — and that proof sets up the next stage: a much, much larger raise.

Get in touch

We are moving
with conviction.

If this is something you want to be part of, in whatever capacity, we would welcome a conversation, and the earlier the better.

contact@featherweights.ai · This page is private — please do not circulate.