featherweights.ai

Nearly everything the AI industry has built since 2017 runs on a single architecture.

The transformer

The transformer arrived in a paper that year, and almost every frontier model since has been the same design made larger. Scaling it up is what made it work. It is also what makes it cost so much. Nine years on, nothing has replaced it.

Who we are

We are a small, passionate team who believe we have discovered an even more substantial breakthrough. We are pursuing commercialization while striking a balance between AI progress, safety and ethics. Featherweights is a simple demonstration of it at a small scale.

What we’ve built

On the evidence so far it shows considerably more capability than the transformer, or any other published approach, while training and running on a small fraction of the compute. We are glad to demonstrate it to anyone who asks.

Confidential

The problem

Every lab has the
same bottleneck.

Ask any frontier lab what limits them and the answer is compute, not ideas or people. Roughly $725B is going into AI infrastructure this year, and what a frontier model needs still climbs four to five times a year. The spend buys headroom and never enough of it.

That one line decides how often a lab can train and what it can afford to serve. It is the number every one of them is trying to bring down, and the reason the biggest names in the field are also the biggest buyers of hardware.

The solution

Our technology saves
99% of that cost.

The advantage has held at every size we have built, and we expect it to widen rather than narrow. It comes from the architecture itself rather than from any technique layered on top of it, and that is the part that compounds as a model grows.

At the spend the industry is running, that is worth billions. The saving is not the main prize. When training costs this little, a lab can attempt models that are out of reach today, and attempt them far more often.

The demo

Meet Tyson.

Tyson is our first chat model, distilled from Qwen3.6-35B-A3B. It punches far above its weight, going toe to toe with the model it learned from while running 10 times faster and costing over 100 times less at this scale.

Tyson gives up a little capability against its teacher, and most deployments would take that trade without thinking twice. It scratches the surface of what the architecture can do. The next model will go a good deal further, on the same economics.

Chat with Tyson

Contact

Get in touch.

If you share our vision, are interested in a partnership, or would like to participate in our pre-seed round, we would like to hear from you.