Sovereign AI — embedded safety, no-code control, zero vendor lock-in
Most institutions can't answer that — the model, the fine-tuning data, and the safety logs all sit on the vendor's servers, not theirs. Cetalabs embeds evaluation and no-code fine-tuning directly in your deployment, so the data, the exports, and the audit trail never leave your roof.
start here
Four ways this goes wrong before anyone notices. Find yours.
why it matters
It shows up on the balance sheet, in public trust, and in the system's own performance — often before anyone calls it a "safety" problem at all.
A pre-deployment evaluation for a mid-size agentic system typically runs RM 15,000–RM 40,000. The remediation, contract renegotiation, and incident-response cost of the same failure caught after launch has run past RM 500,000 in engagements we've seen. The evaluation is the cheap version of finding out. When the model updates come from a foreign vendor, every adaptation requires a new contract or engineering engagement. Embedded fine-tuning and embedded evaluation eliminate that recurring vendor dependency cost.
One visible failure — a leaked bad output, a discriminatory decision, a scam that got through — outweighs a year of correct, unremarkable operation in how an institution gets remembered. Trust doesn't average out; it resets at the worst incident, and recovering it costs more than the incident itself.
Every change request, every language update, every safety patch flows through a vendor you don't control. Sovereign deployment with embedded evaluation and no-code governance removes that bottleneck — and the procurement overhead that comes with it.
RM figures are illustrative ranges drawn from patterns across engagements, not a single cited study — ask us for sector-specific figures (banking, education, health, immigration) for your budget proposal.
A checklist you can hand to procurement or legal before the RFP goes out:
research
Five real failure patterns, each with a runnable check attached — not a hypothetical, a script you can point at your own system.
approach
Terms of service limit the vendor's liability — they don't verify how the system behaves in your languages, your agents, or your jurisdiction. "Trust us" isn't good enough here, and neither is a signature on a contract. Three rules that exist because we've watched each shortcut fail somewhere real, no matter what the T&Cs said.
A safety claim with no attached code is a rumor. Every sovereign deployment includes the script that produces its own audit trail, so you can point it at your own system and get your own answer.
A model that behaves in English and breaks in Manglish didn't pass — it was never tested. So embedded evaluation runs continuously in the languages your citizens actually use, not just once at procurement.
In every red-team run we've done, the failure sat at the seam: agent to agent, agent to human. Nobody's embedded agent "went rogue" — the embedded seam detector catches coordination failures in real-time, before they become contract breaches.
"Compliant with our acceptable-use policy" is a legal shield for the vendor, not evidence the system was tested against your languages, your agents, or your hallucination rate. We measure the behavior directly — the contract language comes after, not instead of, the evidence.
frontiers
Nobody's funding these yet. That's exactly why we think they matter — plain language, no jargon required to see the problem.
When most people turn to the same handful of AI models for advice, opinions, and writing help, the range of ideas they actually encounter quietly narrows — nobody decided this, it just happens. We're studying what that does to how a society thinks, not just what one model outputs. Sovereign fine-tuning allows ministries to produce locally diverse outputs, reducing dependence on a single foreign model's worldview.
epistemic monocultureNearly every AI safety benchmark in existence is written and checked in English, against Western-context examples. Deploy the same system in Bahasa Malaysia or Manglish and the safety behavior may not have come along for the ride — nobody's actually looked. Embedded multilingual evaluation closes this gap continuously — not as a one-time benchmark, but as part of the running system.
global south blind spotMost safety research studies one AI model, alone, in a lab. In practice, damage happens where two systems meet, or where a human trusts an output a beat too fast. That joint is barely studied — and it's where we keep finding the real problems. Embedded multi-agent red-teaming (Alamak Labs) operates at the seam continuously — catching failures the single-model view misses.
seam-over-modelpublic education & training
A one-off audit that leaves the day it's done doesn't practice sovereignty — it just moves the dependency to a different vendor. Every program below runs on a dynamic syllabus: it adapts to what your team already knows, so nobody sits through modules they don't need and nobody drowns in ones they're not ready for.
community
Cetalabs treats outreach the same way it treats research — something you can run and check, not just watch.
Talks and training sessions built around live demonstrations, not slides — an audience watches an AI failure happen, then walks away with a check they can run themselves.
Short, focused sprints that turn a live news story, an invite, or an open question into a runnable mini-experiment within days — kept small on purpose, so the result is checkable rather than sprawling.
products
Nothing here started as a product pitch — each one exists because we hit the problem while doing the research above. Where something isn't fully productized yet, that's stated plainly, not implied.
The Sovereign AI Platform (embedded evaluation + no-code interface) is available as a pilot engagement. BankBench and Substrate modules can be embedded within it or used standalone — ask which applies to your system.
Developed in collaboration with the evaluation, procurement, and sovereign-deployment requirements of banking, foreign-policy, and public-sector engineering teams.
consultation
Public benchmarks don't cover your deployment language, your agent architecture, or your citizens' needs. We don't just evaluate — we deploy embedded AI inside your jurisdiction, with continuous multilingual red-teaming, real-time claim grading, and a no-code interface so your staff control adaptation — not a foreign vendor's engineering team.
Design and deploy embedded AI inside your jurisdiction — with embedded safety harness, data residency guarantees, and fine-tuning rights retained by the ministry.
Continuous embedded red-teaming (VendSafe / Alamak Labs) inside your agent architecture — not a one-time external audit, but a running detector for collusion, coordination failures, and seam breaches.
Configure the embedded no-code interface so your staff adapt outputs and set guardrails. Plus the contract language — audit rights, data residency, update sovereignty, fine-tuning control — for procurement and legal review.
get in touch
Every project above has a runnable version. Tell us which failure sounds like yours, and we'll walk you through the check itself.