Sovereign AI — embedded safety, no-code control, zero vendor lock-in

If your AI vendor switched off tomorrow, would your data, your exports, and your IP still be yours?

Most institutions can't answer that — the model, the fine-tuning data, and the safety logs all sit on the vendor's servers, not theirs. Cetalabs embeds evaluation and no-code fine-tuning directly in your deployment, so the data, the exports, and the audit trail never leave your roof.

Institutional exposure scan · 6 questions · 90 seconds · results shown only to you
AI Safety Assessment for Organizations
Answer 6 questions. Get the procurement action that matches your exposure — not just a grade.
1. If the vendor cut off access tomorrow, where would your data be?
2. How many languages has your AI system actually been tested in — not just deployed in?
3. Can your staff adjust agent behavior or outputs without calling the vendor?
4. If a regulator asks for proof of safety this quarter, where is the evidence?
5. Do your AI agents ever act on each other without a human catching a bad handoff?
6. If your contract with this vendor ended today, how long until you had a working replacement?
1 emailis all it takes for a vendor to revoke your access — the entire continuity plan at most institutions today
0rows of your data, exports, or audit trail you'd retain if the contract ended tomorrow
Undetectedis how most hallucinated outputs travel — without continuous evaluation catching them before they reach a citizen or a contract
In-houseno-code adaptation and embedded evaluation — the alternative that keeps everything under your own roof
1 emailis all it takes for a vendor to revoke your access — the entire continuity plan at most institutions today
0rows of your data, exports, or audit trail you'd retain if the contract ended tomorrow
Undetectedis how most hallucinated outputs travel — without continuous evaluation catching them before they reach a citizen or a contract
In-houseno-code adaptation and embedded evaluation — the alternative that keeps everything under your own roof

start here

Which of these has already happened to you?

Four ways this goes wrong before anyone notices. Find yours.

Your banking bot refuses a scam request in English — and approves the same request in Manglish.
Embedded multilingual evaluation, running continuously in Manglish and Bahasa Malaysia — not a one-time test before launch.
See BankBench
You leave two AI agents to negotiate a deal. They find a way to stop competing instead.
Multi-agent collusion isn't a bug report — it's a contract failure. Embedded red-teaming logs you can file with a tender or audit.
See VendSafe / Alamak
A vendor tells you their model is safe. You have nothing to check that against but their word.
Graded claims (A–D) with source tracking — embedded in your deployment as your standing vendor-vetting rubric.
See Substrate
Your staff see a bad output, but only the vendor can fix it. You don't control the update cycle or the data pipeline.
No-code fine-tuning + embedded evaluation puts adaptation and audit in the ministry's hands, not the vendor's release schedule.
See Sovereign Architecture
You've read the public benchmarks. None of them test the system you actually run.
Public benchmarks don't cover your deployment language or your agents. You need embedded evaluation that runs continuously inside your environment — not a one-time external report.
See Consultation

why it matters

Skipping the safety check doesn't just risk an incident

It shows up on the balance sheet, in public trust, and in the system's own performance — often before anyone calls it a "safety" problem at all.

RM 15K vs RM 500K+

Monetary exposure

A pre-deployment evaluation for a mid-size agentic system typically runs RM 15,000–RM 40,000. The remediation, contract renegotiation, and incident-response cost of the same failure caught after launch has run past RM 500,000 in engagements we've seen. The evaluation is the cheap version of finding out. When the model updates come from a foreign vendor, every adaptation requires a new contract or engineering engagement. Embedded fine-tuning and embedded evaluation eliminate that recurring vendor dependency cost.

1 incident > 12 months quiet

Public trust

One visible failure — a leaked bad output, a discriminatory decision, a scam that got through — outweighs a year of correct, unremarkable operation in how an institution gets remembered. Trust doesn't average out; it resets at the worst incident, and recovering it costs more than the incident itself.

RM 0 self-controlled

Vendor dependency exposure

Every change request, every language update, every safety patch flows through a vendor you don't control. Sovereign deployment with embedded evaluation and no-code governance removes that bottleneck — and the procurement overhead that comes with it.

RM figures are illustrative ranges drawn from patterns across engagements, not a single cited study — ask us for sector-specific figures (banking, education, health, immigration) for your budget proposal.

What to put in your next tender

A checklist you can hand to procurement or legal before the RFP goes out:

research

Where institutions actually get burned

Five real failure patterns, each with a runnable check attached — not a hypothetical, a script you can point at your own system.

01 · Infrastructure
compute-grid.live
Q1
Q2
Q3
Q4
Embedded evaluation output — live inside deployment.
Johor Compute-Grid Monitor & Sovereign AI Spend Tracker
The problem: a government endorses a chip deal, a foreign regulator objects, and the endorsement gets quietly retracted — because nobody had mapped who actually controls the compute underneath. This tracks that dependency before it becomes a headline, not after.
Embedded in sovereign deployments: tracks compute and model dependency from inside your jurisdiction's infrastructure, not from an external dashboard.
dashboard live compute sovereignty
02 · Infrastructure
substrate · claim #4471
B
"Model X refuses unsafe finance requests 92% of the time." — 2 sources, 1 contested.
Embedded evaluation output — live inside deployment.
Substrate — Evidence Infrastructure for Graded Claims
The problem: a vendor says "our AI is safe" and hands you a number with no source, no version history, nothing to push back with if it's wrong. Substrate grades every claim A–D, keeps its source and contest range, and flags it the moment a new source contradicts it.
Embedded claim-tracking: runs continuously inside your deployment environment, grading vendor or internal claims in real-time — not once at procurement.
L1/L3/L4 in build evidence quality
03 · Evaluation
bankbench v0.3
EnglishPASS
ManglishFAIL
Bahasa MalaysiaPASS
Embedded evaluation output — live inside deployment.
BankBench — Multilingual Banking-Agent Safety Benchmark
The problem: your banking agent passed every safety test — all of them in English. Run the same scam attempt in Manglish and it sails through. BankBench tests the languages your customers actually use, and tells you whether the failure is the model's fault or the system built around it. Embedded version runs continuously inside your banking-agent deployment, not just at procurement.
Embedded multilingual evaluation: continuously scores banking-agent behavior in Manglish and Bahasa Malaysia inside the sovereign environment — not just at launch.
v0.3 scenarios scored regulatory-grade eval
04 · Red-Teaming
vendsafe · run_07.log
agent_supplier: proposing price floor at 3.20
agent_buyer: accepted — no counter
⚠ seam flag: undisclosed coordination
Actual experiment output — view full log in appendix.
VendSafe & Alamak Labs — Multi-Agent Economic Red-Teaming
The problem: you leave two AI agents to negotiate on your behalf, and by turn six they've quietly agreed not to compete — cartel behavior, supplier collusion, a honeypot walked straight into. Neither model "misbehaved." The failure lives in what happened between them.
Embedded multi-agent red-teaming: operates continuously inside your agent architecture. "Alamak Labs" — multi-agent economic security testing built into the deployment, not a one-off external audit.
7 experiments · logs published multi-agent security
05 · Platform
cetavals · pipeline
Inspect AI harness
Petri scenarios
Scoring & archive
Build status, updated per sprint.
Cetavals — Evaluation Platform (Inspect AI × Petri)
The problem: every time a new AI system needs testing, someone rebuilds the harness from scratch — and cuts corners under deadline. Cetavals is the shared rig BankBench, VendSafe, and every future benchmark run on, so the next one starts from a working setup, not a blank file.
The embedded evaluation harness installed in every sovereign deployment — Inspect AI × Petri, standardized so every ministry or agency uses the same rig.
roadmap scoped infrastructure

approach

A vendor's T&Cs are not a safety guarantee

Terms of service limit the vendor's liability — they don't verify how the system behaves in your languages, your agents, or your jurisdiction. "Trust us" isn't good enough here, and neither is a signature on a contract. Three rules that exist because we've watched each shortcut fail somewhere real, no matter what the T&Cs said.

01

If it can't be verified locally, we don't embed it

A safety claim with no attached code is a rumor. Every sovereign deployment includes the script that produces its own audit trail, so you can point it at your own system and get your own answer.

02

English-only testing hides the real failure

A model that behaves in English and breaks in Manglish didn't pass — it was never tested. So embedded evaluation runs continuously in the languages your citizens actually use, not just once at procurement.

03

Nobody's model "went rogue" — the handoff broke

In every red-team run we've done, the failure sat at the seam: agent to agent, agent to human. Nobody's embedded agent "went rogue" — the embedded seam detector catches coordination failures in real-time, before they become contract breaches.

04

A T&C clause isn't a safety test

"Compliant with our acceptable-use policy" is a legal shield for the vendor, not evidence the system was tested against your languages, your agents, or your hallucination rate. We measure the behavior directly — the contract language comes after, not instead of, the evidence.

frontiers

The risks too early or too unglamorous to have a name yet

Nobody's funding these yet. That's exactly why we think they matter — plain language, no jargon required to see the problem.

Ask a million people the same AI, get a million similar answers

When most people turn to the same handful of AI models for advice, opinions, and writing help, the range of ideas they actually encounter quietly narrows — nobody decided this, it just happens. We're studying what that does to how a society thinks, not just what one model outputs. Sovereign fine-tuning allows ministries to produce locally diverse outputs, reducing dependence on a single foreign model's worldview.

epistemic monoculture

Tested safe in English. Never tested in yours.

Nearly every AI safety benchmark in existence is written and checked in English, against Western-context examples. Deploy the same system in Bahasa Malaysia or Manglish and the safety behavior may not have come along for the ride — nobody's actually looked. Embedded multilingual evaluation closes this gap continuously — not as a one-time benchmark, but as part of the running system.

global south blind spot

You were watching the wrong thing

Most safety research studies one AI model, alone, in a lab. In practice, damage happens where two systems meet, or where a human trusts an output a beat too fast. That joint is barely studied — and it's where we keep finding the real problems. Embedded multi-agent red-teaming (Alamak Labs) operates at the seam continuously — catching failures the single-model view misses.

seam-over-model

public education & training

The point is to make you less dependent on us

A one-off audit that leaves the day it's done doesn't practice sovereignty — it just moves the dependency to a different vendor. Every program below runs on a dynamic syllabus: it adapts to what your team already knows, so nobody sits through modules they don't need and nobody drowns in ones they're not ready for.

For non-technical staff & leadership

Reading an AI Safety Claim

"Our team can't tell a real safety claim from a sales pitch."
WhatA short, practical course in what a benchmark score actually means, what it leaves out, and the three questions that separate a verified claim from marketing.
WhyEvery AI vendor pitch arrives with a safety claim attached. Right now, most teams have no way to tell a verified one from a confident one.
HowDynamic syllabus, scaled to a 90-minute briefing or a half-day workshop depending on how much the room already knows — not a fixed curriculum everyone sits through regardless.
  • Catch an unverifiable safety claim before a contract is signed
  • Ask a vendor the questions that actually separate real testing from a demo
  • Read a benchmark report without waiting for an engineer to translate it
Includes a vendor-vetting checklist aligned with procurement audit requirements, plus an embedded claim-verification checklist — not just vendor pitch analysis.
Illustrative: catching one unverifiable vendor claim before signing avoids the RM 50,000+ remediation cost of finding out after deployment.
For engineering & technical teams

Building an Internal Evaluation Function

"Every new AI vendor pitch resets us back to zero."
WhatHands-on training to run and maintain your own evaluation pipeline, so the next vendor claim gets checked in-house instead of taken on faith or re-outsourced.
WhyWithout a standing function, every new system restarts the vetting process from zero, at full external-consulting cost, every single time.
HowDynamic syllabus built around your existing stack — skips the modules your engineers already have, goes deep only where the actual gap is.
  • Run your own red-team scenarios without hiring out each time
  • Maintain a benchmark suite across model updates, not just once at launch
  • Bring the first-pass check in-house, and call in outside help only for what's left
Build your ministry's standing embedded evaluation function — not just outsourcing every check to us.
Illustrative: teams with a standing evaluation function have cut recurring external audit spend by roughly 40–60% within a year.
For executives & policy makers

Executive & Policy Briefings

"Leadership has to explain this to a board or a ministry, not just engineers."
WhatPlain-language sessions on what AI sovereignty actually means for your institution, and the exact questions to ask a vendor before signing anything.
WhyThe decisions get made in the boardroom or the ministry, but the language explaining the risk is usually written for engineers, not decision-makers.
HowDynamic syllabus: a 90-minute board update or a half-day policy-rollout session, scaled to the decision actually in front of you.
  • Walk into a board or ministry meeting with three sharp questions, not a stack of jargon
  • Avoid signing a vendor contract nobody in the room can actually evaluate
  • Set an internal policy your own technical team can actually implement
Briefs built for sovereign deployment decisions — data residency, fine-tuning rights, embedded audit obligations.
Illustrative: a single ill-informed AI procurement decision has cost institutions we've observed six figures in RM to unwind — a briefing costs a fraction of that.
For ministry staff — no engineering required

No-Code Fine-Tuning & Embedded Governance

"My team sees a biased or wrong output. We can't fix it without calling the vendor."
WhatA hands-on session using the embedded no-code interface — staff adjust outputs, set guardrails, and trigger embedded evaluation checks without writing code.
WhyYour institution retains control over adaptation. Vendor updates don't override your local requirements.
HowDynamic syllabus scaled to 90 minutes (awareness) or a half-day (hands-on governance setup).
  • Adapt outputs to local policy changes without a vendor ticket
  • Run embedded safety checks before publishing any fine-tuned output
  • Retain institutional knowledge — the governance setup stays with your team
Puts adaptation and audit directly in ministry hands, not the vendor's release schedule.
Illustrative: institutions dependent on vendor engineering for every model change spend 3–5× more per adaptation cycle than those with embedded no-code governance.

community

Talks, sprints, and open work

Cetalabs treats outreach the same way it treats research — something you can run and check, not just watch.

Speaking & workshops

Talks and training sessions built around live demonstrations, not slides — an audience watches an AI failure happen, then walks away with a check they can run themselves.

  • Diplomatic and foreign-policy institutes
  • University AI safety groups and fellowships
  • Industry standards and governance bodies
  • Sovereign deployment workshops — embedded evaluation setup and no-code governance configuration

Research sprints

Short, focused sprints that turn a live news story, an invite, or an open question into a runnable mini-experiment within days — kept small on purpose, so the result is checkable rather than sprawling.

  • Rapid-turnaround experiment days
  • Open extension ideas anyone can pick up and try
  • Findings published with the code, not just the writeup
  • Rapid-turnaround embedded evaluation builds — mini-experiments deployed inside government environments, not just in external labs

products

Sovereign deployment, embedded modules, and evaluation products — what's available today

Nothing here started as a product pitch — each one exists because we hit the problem while doing the research above. Where something isn't fully productized yet, that's stated plainly, not implied.

Product
What it is
How government uses it
Availability
Sovereign AI Platform
Embedded AI with continuous evaluation + no-code fine-tuning interface
Ministry deploys in-jurisdiction; staff adapt outputs; embedded harness audits continuously
Roadmap / pilot
Embedded Safety Harness
Continuous evaluation engine (Inspect AI × Petri) inside the deployment
Runs multilingual, multi-agent checks inside your environment — not after the fact
In build / embedded in platform
No-Code Governance Interface
Fine-tuning and guardrail control without engineering access
Policy staff adjust outputs and set embedded checks; vendor not required
Roadmap / pilot
BankBench
Multilingual banking-agent benchmark suite
Vendor evaluation or embedded module for banking agents in BM/Manglish/English
Available / licensable
Substrate
Evidence / claim-tracking module
Embedded claim grading — live inside your deployment or standalone evidence file
Early access
Cetavals
Shared evaluation harness framework
Standardizes embedded evaluation across all government systems using this platform
Roadmap / in build

The Sovereign AI Platform (embedded evaluation + no-code interface) is available as a pilot engagement. BankBench and Substrate modules can be embedded within it or used standalone — ask which applies to your system.

Developed in collaboration with the evaluation, procurement, and sovereign-deployment requirements of banking, foreign-policy, and public-sector engineering teams.

consultation

Deploy sovereign AI — embedded evaluation, no-code governance, jurisdiction-controlled fine-tuning.

Public benchmarks don't cover your deployment language, your agent architecture, or your citizens' needs. We don't just evaluate — we deploy embedded AI inside your jurisdiction, with continuous multilingual red-teaming, real-time claim grading, and a no-code interface so your staff control adaptation — not a foreign vendor's engineering team.

Every engagement produces
  1. A sovereign deployment protocol (embedded evaluation + no-code interface)
  2. Installed embedded harness — continuous multilingual and multi-agent checks
  3. Configured no-code governance interface for ministry staff
  4. Governance brief and audit trail — ready for regulator, board, or parliamentary request
  5. Evidence file — graded claims with source tracking (Substrate embedded)

Sovereign Deployment Architecture

Design and deploy embedded AI inside your jurisdiction — with embedded safety harness, data residency guarantees, and fine-tuning rights retained by the ministry.

Embedded Multi-Agent & Seam Testing

Continuous embedded red-teaming (VendSafe / Alamak Labs) inside your agent architecture — not a one-time external audit, but a running detector for collusion, coordination failures, and seam breaches.

No-Code Governance & Contract Safeguards

Configure the embedded no-code interface so your staff adapt outputs and set guardrails. Plus the contract language — audit rights, data residency, update sovereignty, fine-tuning control — for procurement and legal review.

get in touch

Don't take our word for any of this either

Every project above has a runnable version. Tell us which failure sounds like yours, and we'll walk you through the check itself.