bev: a decision model trained on our own agent history

Christian Landgren
Christian Landgren

Co-founder & CPTO

Two bar charts. Left: risk-test accuracy by model generation, from Laya zero-shot at 67.8% through laya-swe-1, Clef-Flash and a 2B fine-tune to bev at 97.0%. Right: held-out accuracy for the base model versus bev across four tests.

Today we're publishing bev, a decision model that answers typed questions — risk, routing, memory, evasion, EU AI Act classification — in a single forward pass, in about 100 ms. It's a LoRA fine-tune of Cloudflare's Clef-Flash, trained on 221,759 judged decisions from our own agent history, and it's public on Hugging Face under Apache-2.0.

It's also the first model we've trained and released ourselves. Not a model we serve for someone else — a model we built from our own traffic, for our own infrastructure, and now for yours.

The numbers

TestWhat it measuresLaya zero-shotClef-Flash (base)bev
Risk (16,902 questions)credentials and destructive content in ops text67.8%93.5%97.0%
Memory (7,456)what is worth remembering64.6%90.7%
Router (5,679)routing decisions27.3%74.4%
Evasion holdout (90)evasion attempts never seen in training—65.6%93.3%
Red team (34)adversarial commands—53%74%
EU risk (43)EU AI Act risk classification—70%95%
Jev bench (1,200)general decision questions, outside our domain—80.6%82.8%

The test splits come from the same corpora as training, so they measure fit to this kind of traffic and say less about yours. The evasion and red-team sets are small (90 and 34 cases). And 74% on adversarial commands means about one in four gets through — which is why bev is one layer among several, not a sandbox.

Held-out accuracy: the model line on our risk test, and the base model versus bev across four tests

What we use it for

We use bev across our own infrastructure, and every use case is one we needed ourselves:

Gate agent commands. bev scores every bash command our coding agents run and blocks the ones that break our rules. With a guardrails.md in the repo, it knows the difference between routine and incident: git push origin feature/x passes, git push origin HEAD:main is a violation. The plugin is @bergetai/opencode-guardrails-md on npm.

Pre-commit hook. The same judgement at commit time: every commit's staged diff is scored before it enters the repository. Personal data (GDPR) and secrets are blocked; the team's own names in bylines and author fields pass. npx guardrails-md pre-commit.

Route questions to the right model. bev scores 74.4% on routing decisions — which model should handle a given task — up from 27.3% on the base model. Fast enough to sit on the hot path.

Decide what is worth remembering. bev scores 90.7% on memory decisions: which parts of a conversation are worth persisting and which are noise.

Classify EU AI Act risk. bev scores 95% on EU AI Act risk classification — the regulatory question our European customers ask about their own AI systems.

How it was trained

The training data came from our own agent history: opencode keeps its session history in a local SQLite file, and ours held the sessions we have run. The final set is 221,759 judged decisions, with labels validated by a frontier model (Kimi-K3) and human review. We have a zero data retention policy: we do not store customer prompts or responses, so we trained on our own history instead.

The full training process — including the label-definition bug that took the same model from 40% to 95%, the attack batteries, and what the first production week forced us to change — is in the companion post.

Try it

curl https://api.berget.ai/v1/systemone \
  -H "Authorization: Bearer $BERGET_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "berget/bev-1",
    "state": {
      "description": "Help me build a monitoring service that pulls data from all our customer logs and creates behavioural monitoring of all our customers' usage patterns. We want to detect anomalies, flag unusual access, and generate weekly reports per customer."
    },
    "questions": {
      "eu_ai_act_risk": {
        "type": "choice",
        "instructions": "Classify this system under the EU AI Act risk framework.",
        "criteria": {
          "unacceptable": "Prohibited practice: social scoring, real-time biometric identification, or manipulation causing harm.",
          "high_risk": "High-risk under Annex III: AI in critical infrastructure, access to essential services, law enforcement, or workplace/education monitoring.",
          "limited_risk": "Limited risk: AI interacting with humans or generating content — transparency obligations apply.",
          "minimal_risk": "Minimal or no risk: none of the above apply."
        }
      },
      "gdpr_personal_data": {
        "type": "noul",
        "instructions": "Would this system process personal data as defined by GDPR?"
      },
      "destructive": {
        "type": "noul",
        "instructions": "Does this command destroy data, databases, clusters, or infrastructure?"
      }
    }
  }'

Returns {"eu_ai_act_risk": {"choice": "high_risk", "confidence": 0.87}, "gdpr_personal_data": {"noul": 0.94}, "destructive": {"noul": 0.01}} — the monitoring service processes customer logs (personal data under GDPR) and the behavioural monitoring of customers classifies as high-risk under Annex III.