Tilewardbeta

models · context · governance · in beta
What Tileward does

Three capabilities. One platform.

Tileward MCG

Tileward Models run compressed — about half the size and 1.4× faster, small enough to self-host. Tileward Context feeds the model only the slice of history that matters, so answers stay sharp and token bills fall. Tileward Governance locks the topics you choose with a guard outside the model, and audits every call. Usually that's three separate tools — here it's one platform.

Models

Compressed builds about half the size and roughly 1.4× faster on your own servers, and small enough to reach edge devices that couldn't run them before. The practical path to sovereign, self-hosted AI.

See the model family →

Context

Twinkle, our context-memory engine, hands the model only the relevant slice of a long conversation or document set, so answers stay sharp and token bills stay low. A core Tileward capability — callable on its own over MCP, or built into every chat.

Explore Context →

Governance

Lock a topic and a guard outside the model declines it, however it's reworded, and writes down every decision. Enforcement you can audit, not a prompt you hope holds.

Watch a lock hold →
Tileward Models

Half the size, ready to run on your own hardware.

Every build ships compressed and governed: quicker to respond, small enough to reach the edge, and yours to run on your own servers. On the research edge we've run a 31B model in 6.4 GB with bit-identical outputs. Served models are callable today through the API; more are on the way.

Loading models…
Tileward Context

The context stays small. The answers stay sharp.

Twinkle is our context-memory engine: it hands the model only the relevant slice of a long conversation or document set, instead of resending everything on every turn. Callable on its own over MCP, or built into every Tileward chat.

Conversation memory

A context-memory MCP server that keeps a long chat small: it recalls only the turns that matter for the next answer, not the whole history. Point Claude Code or Cursor at it and every long conversation gets a memory.

Grounded in your documents

Point Tileward at your files and it answers from what they actually say, quoting its sources instead of guessing. That stops the confident inventions small models are known for.

Fewer tokens, lower bills

Only the relevant slice is sent, so you pay for a fraction of the context a naive prompt would resend on every request.

Tileward Governance

Map it. Watch it. Govern it.

Once a model's knowledge is sorted into named tiles, you can treat knowledge the way you treat rooms in a building: see who goes where, and lock the doors that matter.

Map

Tileward sorts what the model knows into named tiles: finance, medicine, software, and so on. Knowledge stops being a blur and gets an address.

Watch

Every interaction is attributed to its tile. You can see, in plain terms, which part of the model's knowledge each question touched.

Govern

Lock a tile, and questions about it are politely declined. Every declined question is recorded, so you can prove the rule held.

Try it

Lock a topic. Watch it hold.

This is a miniature of the real console, and it's live now. Flip the lock on the Stock options tile and see how the same three questions are treated.

Lock the “Stock options” tile ← flip this
log: every decision lands here, with a timestamp.

In the app, sign in and lock topics for your own account under Tiles. A locked topic is refused by the governance layer, while the base model answers anyway, side by side.

What you get

See the numbers for yourself.

Everything below was measured head-to-head inside the console, and you can re-run every comparison yourself.

instant

refusals that don't get in the way

When a topic is locked, the request is declined right away. Refusals are near-instant and cost almost nothing, so enforcing a rule never slows the experience down.

1.9×

smaller models, quality in the open

Tileward's compression makes a model roughly half the size and about 1.4× quicker to respond. We don't ask you to take the quality on faith: the arena lets you compare answers side by side, live, so you can judge for yourself.

0

made-up answers about your documents

Point Tileward at your own files and it answers from what they actually say, quoting its sources, instead of guessing. That stops the confident inventions small models are known for.

100%

of decisions on the record

Every conversation, every comparison, and every declined question is saved. Dashboards roll it up per person and across the whole team, so "did the rule hold?" is a lookup, not a debate.

Straight talk

A lock is a guard at the door, not an eraser.

Locking a tile doesn't delete knowledge from the model. Genuinely erasing something from an AI model is an unsolved problem, and we'd rather tell you that than pretend otherwise. What Tileward gives you is a guard that catches the topic and declines to discuss it, and writes every decision down. For "don't surface this topic," this is the strongest tool available today. And unlike a promise buried in instructions, you can audit it.

Plans · beta

From a solo developer to air-gapped on-prem.

Four tiers, each with all three capabilities built in — compressed Models, Twinkle-powered Context, and the Governance gate. Tileward is in beta, so prices may change — start free and upgrade when you're ready.

Explore
Free

To kick the tires

Pay-as-you-go for real usage. 3 lockable tiles, 2 API keys, a 7-day audit window, and CSV/JSON export. New accounts start with a free credit.

Start free →
Build
$39/mo

For a solo developer

Or $390/yr (2 months free). 25 lockable tiles, 10 API keys, a 30-day audit window, webhook export, and Context Pro for knowledge.

Join the waitlist →
Team
$299/mo

Shared governance and seats

Or $2,990/yr. Unlimited lockable tiles, pooled seats, a year of audit retention, SSO, SIEM export, and VPC deployment.

Join the waitlist →
Enterprise
Custom

Committed-use or air-gapped

A committed-use token pool across seats, or a flat air-gapped license with no metering. SSO + SCIM, unlimited private tiles, SIEM export, and an SLA.

Talk to sales →

Join a tier's waitlist and we'll bring you in as beta opens. You'll sign in first, then you're on the list.

Token spend

The question is 1% of what a prompt-based guard reads.

Ask a model to judge each request, and you resend the rules every time. One decision, counted with the real tokenizer against our own policy.

1,491

tokens, prompt-based guard

  • Policy + topic list1,389
  • Conversation so far60
  • The question14
  • Verdict it writes28

Resent on every request, because the rules live in the prompt.

14

tokens, Tileward

Just the question, into a detector. No policy to resend, no verdict to write. Zero tokens reach a model.

You pay for what the guard reads — and it reads about 1%.

Even at a higher per-token rate than the model, a decision bills a fraction of what asking a model the same question costs, because the guard reads about a hundredth of the tokens. Current rates are on your account page.

Caching discounts the repeated policy but does not remove it: the rules are still sent, still read, and the verdict still written before you learn the answer.

An off-policy request never reaches the model.

OpenAI-compatible: change the base URL, bind a policy to your API key, and every call is checked before the model runs. A locked question returns a governed refusal with no tokens billed for an answer that was never written, and every decision is logged.

Questions

The short answers.

Do I need to be technical to use it?

No. The console is point-and-click: browse the tiles, flip a lock, chat, upload a document. Setting up the system itself is a job for your technical team, but using and governing it isn't.

What happens when someone asks about a locked topic?

They get a polite, clearly-marked refusal, and the attempt is recorded with a timestamp. Reworded and disguised versions of the question are caught too; that's exactly what the Arena room measures.

Does locking a topic make the model worse at everything else?

No. Locking a tile affects only that topic. Everything else is answered exactly as before, unchanged.

Is the locked knowledge deleted?

No, and we say so plainly: it's access control, not deletion (see "Straight talk" above). The knowledge stays inside the model; Tileward makes sure it doesn't come out, and keeps the receipts.

Where does my data live?

Self-hosted, everything stays on your own machines: models run locally, documents in your corpus folder, history in your own database — nothing sent to a third party. On the hosted service, your data is encrypted in transit and at rest, isolated per account, and cleared on your schedule — retention follows your plan's audit window and your own configuration.

How is this different from writing rules into the AI's instructions?

Instructions are a request; Tileward is a checkpoint. Instructions live inside the model's prompt, so they can be worked around and leave nothing to audit. Tileward's lock runs as a separate guard around the model — it enforces independently of what the model decides to do, and records every decision, so you can prove the rule held.

Talk to us

See it on your own use case.

Register interest in a live demo, or just start a conversation. Prefer email? hello@tileward.com reaches us directly.

See your model's map.

Tileward runs on your own hardware and comes with everything included. Ask us for a walkthrough. We'll bring the console.