faxl

Sovereign AI

navigate the frontier

If you buy billions of tokens a day, it is time to manage your tokenomics directly: rent tin in your own VPC, run your own models behind one proxy, and own the cache.

Book a meeting

Why faxl

Your cache protects someone else's margin

Every prefill cache in the market today — Foundry, Bedrock, and the self-host stacks — lives inside the inference provider and exists to maximise the seller's throughput per GPU. You are a token consumer. You cannot see that cache, control it, keep it, or move it.

The incentives are not aligned with yours. When the cache is invalidated you pay to process the same text again, and that is revenue to the seller. At billions of tokens a day, the prefill an opaque cache silently re-charges is an eight-figure line item nobody in your organisation can defend, because nobody can see it.

faxl is an OpenAI-compatible proxy in front of inference you control. Existing clients change one URL.

Model control

Be the registry, not the tenant

Your organisation already refused to build on public package registries it did not control — that is why you run Artifactory or Nexus. Models are the same problem, one layer up. Weights pulled from a public hub on trust, a different revision on every node, and no record of which model answered which question.

faxl makes the proxy the gate. Register a model once, serve it from your VPC, pin the revisions that are allowed, and every request in the organisation passes one place that knows what ran, for whom, and what it cost.

Model registry
Approved models and revisions, served from your estate — not pulled from a public hub at run time.
One address
An OpenAI-compatible endpoint for the whole enterprise. Clients change a URL, nothing else.
Cache prefill control
You set what is warmed and kept per model, rather than discovering the vendor's policy from your bill.

Enterprise controls

What a registry gives you, applied to inference

faxl is to prefill state and models what an artifact repository is to builds. These are the controls a platform team expects before it will put shared infrastructure into production.

Audit trail
Every request logged durably — prompt hash, prefix-hash array, lookup probes, cache decision, outcome — in deterministic order, and replayable.
Per-tenant telemetry
A canonical namespace over tenant, model, revision, backend, dtype and config is stamped on every record, so every request in the estate is attributable. It also stops state crossing a model or engine boundary, where reuse would be wrong.
Verified before use
Exact prefix-hash plus token-for-token confirmation before any resume. Sellers load on trust; faxl verifies.
Retention policy
You choose what persists and for how long. Correctness never depends on the choice.
Cost attribution
Tokens served, tokens skipped and bytes stored, per department, per model, per application. Invisible re-prefill becomes a line item someone owns.
Internal cross-charging
A monthly statement per team for AI spend on shared infrastructure, so each department pays for what it used.
Adoption analytics
Who is using AI, how much, and whether it is growing. Measured from the same record, not surveyed.
Survives churn
State persists off-GPU and restores across restart, deallocation and cross-node routing — the conditions where a vendor cache evaporates.
Long-gap workloads
A warmed context survives between runs. A frozen corpus hit once a morning, or a batch job that wakes and sleeps, keeps its cache across the idle gap.
Portability
The cache is an off-GPU blob keyed by namespace, not tied to a live process. Move a workload between clouds or on-prem and the warm state moves with it.
Air-gapped operation
No phone-home, no licence server, no managed dependency. Lift the whole thing onto sovereign bare metal and every control comes with it.

Roadmap

The proxy is the surface of your AI service

A shared AI team inside a large organisation is trying to offer one secure, sovereign AI capability to thousands of staff and their agents. Everything that request touches — which model, which documents, what is remembered, what it cost — is decided at the proxy. That makes the proxy the service, not a component of it.

The direction of travel is a faxl interface your staff can use directly: the ergonomics people expect from a public chat assistant, serving your own models, with retrieval against your own corpora rather than an external service. Same gate, same audit, same attribution — and nothing leaves the estate.

What ships today is the proxy and the cache. Ask us where any capability above stands.

Measured

What is proven today

  • Repeated prompt on Kimi Linear 48B: first token 5,056 ms → 246 ms, byte-identical output.
  • Proven on the hybrid recurrent architectures — Granite 4.0-h, Kimi Linear 48B, Qwen3-Next-80B — the designs that break ordinary prefix caching. Published analysis of Kimi K3 infers the same architecture.
  • Real agent traffic through a warm store: ~70% of prompt tokens skipped.
  • Exact re-run of a 40-task workload: 92% skipped.
  • 570/570 adversarial reuse decisions correct; no wrong or stale state served.
  • The NVIDIA/vLLM port passes the same byte-identity bars on GPU.

Results on your estate depend on your models, your traffic, and how much your prompts repeat. Bring your traffic profile and we will measure it.

Limits

Where faxl does not help

faxl needs access to the inference engine, so it works with self-hosted vLLM on NVIDIA and MLX on Apple silicon. It cannot attach to a pure managed API — which is precisely the case where you are most captive to the seller's margin. The pain is real even where we cannot yet serve it, and the move that answers it is bringing the workload onto infrastructure you control.

Several of the controls above are specified and in build rather than shipped. Ask us which, and we will tell you exactly where each one stands.

Talk to us

Bring your traffic profile and model mix. We will show you what a warm cache takes off the bill, and what it would take to run it on your own tin.

Book a meeting

Andrew Morgan, Gamakon Ltd — andrew@gamakon.ai