Research lab · Model efficiency

We are not trying to make AI
smarter
but far more efficient.

Artificial intelligence currently advances by spending more: more parameters, more memory, more compute, more energy. Symbioz explores the other direction — getting the same for far less. Not as a finishing touch, but as the whole of our research.

0 of memory per token, today
0 less than published references
0 consumer GPU per experiment
0 points of quality traded away
Efficiency Inference memory Bytes per token Long context Silent reasoning Quantization Hardware frugality Negative results
Why we exist

Scaling up has a price

For a decade, progress in artificial intelligence has followed a single recipe: make it bigger. More parameters, more data, more silicon. It works — it has produced remarkable systems. But it also carries a cost that rarely gets examined.

A model that is twice as frugal delivers the same service for half the resources — on every single request, forever. At the scale AI is used today, a gain in efficiency is worth more than one more point of benchmark performance.

That is the direction Symbioz works in, and the only one. We are not trying to beat anyone on a capability leaderboard. We are trying to fit the same thing into far less.

Energy

Every byte a model does not have to keep is memory that is never manufactured and power that is never drawn — on every request, indefinitely.

Access

Training and serving a very large model is the preserve of a handful of organizations. What becomes efficient becomes reachable again for small teams.

Autonomy

A model that fits on an ordinary machine can run at home, offline, without handing your data to anyone.

Our work

Six places to win efficiency

The same question every time: can we get the same result for less? What follows is what we are after and why — the architectural details stay in the lab.

Useful reach

A model that "accepts" 128,000 tokens does not thereby use them. Advertising a window costs nothing; actually using it costs a great deal. We measure the reach you pay for.

  • Where the model genuinely breaks down
  • Telling the advertised window from the useful reach
  • Going further without paying more

Silent reasoning

Making a model think currently costs thousands of tokens, written and then thrown away. What if it formed a picture of what comes next before writing anything at all?

  • Anticipating rather than guessing word by word
  • Reasoning that is not billed by the token
  • Exploratory work, in progress

Just enough precision

A model computes with far more decimal places than it needs, and each one costs memory and energy. We step down one level at a time, measuring what is lost at each.

  • 16 bits, then 8, then 4
  • What degrades first, and on which texts
  • The real gain, not the theoretical one

The part that works

In a trained model, not everything earns its keep. Some parameters no longer contribute anything yet cost exactly as much as the rest. We want to count them, then stop paying for them.

  • Measuring the genuinely active share
  • What can be removed without changing the output
  • Designing without that dead weight from the start

The hardware constraint

Everything we publish is trained and measured on consumer GPUs, one at a time. This is not a limitation we endure: it is the constraint that forces us to find the saving.

  • One GPU per experiment, not a data center
  • Results a small lab can reproduce
  • Efficiency as the starting point, not a late fix
What it comes down to

Where efficiency is actually measured

Not in the parameter count, the number everyone puts on the box. In the bytes a model must keep for every token it reads. In four steps.

01

The model reads

Every token read leaves a trace the model has to carry for the rest of its answer.

02

The trace piles up

The longer the text, the larger that memory grows — token after token, never emptying.

03

The wall

Once the available memory runs out, that trace sets the real limit: it caps the context and drives the bill.

04

Efficiency

Compressing that trace to 90 bytes per token — then checking, text by text, that nothing was lost on the way.

0 of memory per token in our model
0 for an open reference architecture
0 consumer GPU per experiment
0 results announced without a preregistered protocol
How we work

Deciding before looking

Chasing efficiency makes it unusually easy to fool yourself: the gains you are hunting are small, and a tiny difference looks a lot like noise. So we took the opposite approach to intuition — acceptance criteria are written down, dated and frozen before each experiment starts.

What we measure afterwards is not open to interpretation. An appealing idea that lands in the grey zone is dropped — even if we liked it, even if it cost us weeks. Most of our leads end that way, and that is simply what a lab looks like from the inside.

A saving only counts as a gain if quality has not moved. It is the least spectacular half of our work, and by far the longest.

Preregistration

Acceptance thresholds, stopping rules and analysis strata are fixed before a single experiment is launched.

Paired comparison

Two architectures are compared on identical texts, with a confidence interval, and the result is broken down by type of passage. An average gain that hides a local loss is not a gain.

Negative results

Abandoned leads are documented as carefully as the successes. Knowing which saving does not work is worth something.

Transparency

What we share, what we keep

Better said plainly than left to be guessed.

What we publish

The questions we ask, the reasons we chose them, the savings we measure and the results we get — including the ones that prove us wrong. Formal publications will follow once our large model is finished; until then, this site serves as the record.

What we keep

Architectural detail, training recipes and internal mechanisms stay in the lab. We communicate the what and the why, not the how.

What we will not do

No in-house leaderboard cut to flatter us, no demonstration picked after the fact, no saving announced without stating what it cost in quality.

The lab

Small, independent, and slow on purpose

Symbioz is an independent research lab based in France. We have nothing to sell: this site exists to report on our work, not to turn a visitor into a customer.

We work on deliberately compact models. Not for want of anything better, but because it is the right scale at which to study efficiency: an idea that adds nothing shows up immediately, where a very large model would absorb it unnoticed.

Results obtained at that scale are then re-examined before any generalization. We would rather have a narrow, solid conclusion than a broad promise.

One single question

Getting the same for less. Anything that does not serve that question leaves the research programme.

Independence

No outside funding to satisfy, and therefore no reason to force a result or announce it too early.

Traceability

Every published result points back to a dated experiment, its protocol and its reference data, all archived.

Contact

Working on something close?

We read everything that touches model efficiency: inference-memory compression, quantization, honest evaluation. A remark, an objection, a result that contradicts ours — write to us.