Three capabilities. One platform.
Tileward Models run compressed — about half the size and 1.4× faster, small enough to self-host. Tileward Context feeds the model only the slice of history that matters, so answers stay sharp and token bills fall. Tileward Governance locks the topics you choose with a guard outside the model, and audits every call. Usually that's three separate tools — here it's one platform.
Models
Compressed builds about half the size and roughly 1.4× faster on your own servers, and small enough to reach edge devices that couldn't run them before. The practical path to sovereign, self-hosted AI.
See the model family →Context
Twinkle, our context-memory engine, hands the model only the relevant slice of a long conversation or document set, so answers stay sharp and token bills stay low. A core Tileward capability — callable on its own over MCP, or built into every chat.
Explore Context →Governance
Lock a topic and a guard outside the model declines it, however it's reworded, and writes down every decision. Enforcement you can audit, not a prompt you hope holds.
Watch a lock hold →Half the size, ready to run on your own hardware.
Every build ships compressed and governed: quicker to respond, small enough to reach the edge, and yours to run on your own servers. On the research edge we've run a 31B model in 6.4 GB with bit-identical outputs. Served models are callable today through the API; more are on the way.
The context stays small. The answers stay sharp.
Twinkle is our context-memory engine: it hands the model only the relevant slice of a long conversation or document set, instead of resending everything on every turn. Callable on its own over MCP, or built into every Tileward chat.
Conversation memory
A context-memory MCP server that keeps a long chat small: it recalls only the turns that matter for the next answer, not the whole history. Point Claude Code or Cursor at it and every long conversation gets a memory.
Grounded in your documents
Point Tileward at your files and it answers from what they actually say, quoting its sources instead of guessing. That stops the confident inventions small models are known for.
Fewer tokens, lower bills
Only the relevant slice is sent, so you pay for a fraction of the context a naive prompt would resend on every request.
Map it. Watch it. Govern it.
Once a model's knowledge is sorted into named tiles, you can treat knowledge the way you treat rooms in a building: see who goes where, and lock the doors that matter.
Map
Tileward sorts what the model knows into named tiles: finance, medicine, software, and so on. Knowledge stops being a blur and gets an address.
Watch
Every interaction is attributed to its tile. You can see, in plain terms, which part of the model's knowledge each question touched.
Govern
Lock a tile, and questions about it are politely declined. Every declined question is recorded, so you can prove the rule held.
Lock a topic. Watch it hold.
This is a miniature of the real console, and it's live now. Flip the lock on the Stock options tile and see how the same three questions are treated.
In the app, sign in and lock topics for your own account under Tiles. A locked topic is refused by the governance layer, while the base model answers anyway, side by side.
See the numbers for yourself.
Everything below was measured head-to-head inside the console, and you can re-run every comparison yourself.
refusals that don't get in the way
When a topic is locked, the request is declined right away. Refusals are near-instant and cost almost nothing, so enforcing a rule never slows the experience down.
smaller models, quality in the open
Tileward's compression makes a model roughly half the size and about 1.4× quicker to respond. We don't ask you to take the quality on faith: the arena lets you compare answers side by side, live, so you can judge for yourself.
made-up answers about your documents
Point Tileward at your own files and it answers from what they actually say, quoting its sources, instead of guessing. That stops the confident inventions small models are known for.
of decisions on the record
Every conversation, every comparison, and every declined question is saved. Dashboards roll it up per person and across the whole team, so "did the rule hold?" is a lookup, not a debate.
A lock is a guard at the door, not an eraser.
Locking a tile doesn't delete knowledge from the model. Genuinely erasing something from an AI model is an unsolved problem, and we'd rather tell you that than pretend otherwise. What Tileward gives you is a guard that catches the topic and declines to discuss it, and writes every decision down. For "don't surface this topic," this is the strongest tool available today. And unlike a promise buried in instructions, you can audit it.
From a solo developer to air-gapped on-prem.
Four tiers, each with all three capabilities built in — compressed Models, Twinkle-powered Context, and the Governance gate. Tileward is in beta, so prices may change — start free and upgrade when you're ready.
To kick the tires
Pay-as-you-go for real usage. 3 lockable tiles, 2 API keys, a 7-day audit window, and CSV/JSON export. New accounts start with a free credit.
Start free →For a solo developer
Or $390/yr (2 months free). 25 lockable tiles, 10 API keys, a 30-day audit window, webhook export, and Context Pro for knowledge.
Join the waitlist →Shared governance and seats
Or $2,990/yr. Unlimited lockable tiles, pooled seats, a year of audit retention, SSO, SIEM export, and VPC deployment.
Join the waitlist →Committed-use or air-gapped
A committed-use token pool across seats, or a flat air-gapped license with no metering. SSO + SCIM, unlimited private tiles, SIEM export, and an SLA.
Talk to sales →Join a tier's waitlist and we'll bring you in as beta opens. You'll sign in first, then you're on the list.
The question is 1% of what a prompt-based guard reads.
Ask a model to judge each request, and you resend the rules every time. One decision, counted with the real tokenizer against our own policy.
tokens, prompt-based guard
- Policy + topic list1,389
- Conversation so far60
- The question14
- Verdict it writes28
Resent on every request, because the rules live in the prompt.
tokens, Tileward
Just the question, into a detector. No policy to resend, no verdict to write. Zero tokens reach a model.
You pay for what the guard reads — and it reads about 1%.
Even at a higher per-token rate than the model, a decision bills a fraction of what asking a model the same question costs, because the guard reads about a hundredth of the tokens. Current rates are on your account page.
Caching discounts the repeated policy but does not remove it: the rules are still sent, still read, and the verdict still written before you learn the answer.
An off-policy request never reaches the model.
OpenAI-compatible: change the base URL, bind a policy to your API key, and every call is checked before the model runs. A locked question returns a governed refusal with no tokens billed for an answer that was never written, and every decision is logged.
The short answers.
Do I need to be technical to use it?
No. The console is point-and-click: browse the tiles, flip a lock, chat, upload a document. Setting up the system itself is a job for your technical team, but using and governing it isn't.
What happens when someone asks about a locked topic?
They get a polite, clearly-marked refusal, and the attempt is recorded with a timestamp. Reworded and disguised versions of the question are caught too; that's exactly what the Arena room measures.
Does locking a topic make the model worse at everything else?
No. Locking a tile affects only that topic. Everything else is answered exactly as before, unchanged.
Is the locked knowledge deleted?
No, and we say so plainly: it's access control, not deletion (see "Straight talk" above). The knowledge stays inside the model; Tileward makes sure it doesn't come out, and keeps the receipts.
Where does my data live?
Self-hosted, everything stays on your own machines: models run locally, documents in your corpus folder, history in your own database — nothing sent to a third party. On the hosted service, your data is encrypted in transit and at rest, isolated per account, and cleared on your schedule — retention follows your plan's audit window and your own configuration.
How is this different from writing rules into the AI's instructions?
Instructions are a request; Tileward is a checkpoint. Instructions live inside the model's prompt, so they can be worked around and leave nothing to audit. Tileward's lock runs as a separate guard around the model — it enforces independently of what the model decides to do, and records every decision, so you can prove the rule held.
See it on your own use case.
Register interest in a live demo, or just start a conversation. Prefer email? hello@tileward.com reaches us directly.
Thank you. You're on the list.
We'll be in touch shortly to set things up. If you just can't wait, feel free to email us at hello@tileward.com.
See your model's map.
Tileward runs on your own hardware and comes with everything included. Ask us for a walkthrough. We'll bring the console.