Platform · 04 Improve
Detection quality decays: sources change, environments drift, and the rule that was precise in March is noise by August. Qvasir feeds real outcomes back into the rules that produced them, so tuning stops being a quarterly project and becomes a property of the system.
Calibration
Tuning normally starts with a blank tab and a decision to go and measure something. Here the measuring has already happened: save a generated rule and its calibration against your own history is dispatched without being asked, so what you arrive at is a false-positive surface rather than an empty panel and an intention.
Candidate values to exclude are not one model's opinion. Each is judged three times and aggregated, and you see the level of agreement alongside the proposal — so a unanimous call and a narrow one do not look identical on the screen.
A proposal that would whitelist the rule's own trigger value is rejected in code, whatever the model concluded. That is the one exclusion that would silently neuter the rule, so it is not left to judgement.
Past a threshold of distinct values, the platform advises against excluding at all — deterministically. A field with hundreds of values is telling you the rule is wrong, not that you need a longer allowlist.
The guardrail that matters
This is the part worth reading carefully, because it is a design decision rather than a feature. Wherever a model touches a live detection, what it is permitted to do is constrained by code — not by an instruction in a prompt that a cleverly-worded log line might talk it out of.
The AI reviewing your calibration results can demote a value to "review carefully" or attach a warning. It structurally cannot approve a value, raise a confidence score, or change any number — the merge is an add-only operation. And if a merge ever ended up relaxing the deterministic verdict, the entire AI contribution is discarded rather than trusted. It can only ever narrow.
When it refines a rule from a tuning suggestion, it may change the detection logic and the tunables. Anything it returns outside that scope is reverted before you ever see it — so "fix the false positives" cannot quietly become a rewritten severity, a changed source binding, or a new set of ATT&CK tags.
Hand it a tuning suggestion and it can come back rejecting it, per suggestion, with a reason. An assistant that implements every instruction including the bad ones is not much of a safeguard.
Calibration data contains values an attacker can influence — a hostname, a filename, a command line. Those are fenced inside a delimited block that is regenerated per run, and the model reasons over aggregates rather than the raw stream. Prompt injection through your own telemetry is a real attack on a platform like this one.
The pattern underneath all four is the same: the AI proposes, deterministic code decides what a proposal is even allowed to be, and a person applies the ones that reach a live system. Every applied change lands in the rule's own version history.
When tuning is not the answer
When a group of rules is noisy for the same underlying reason, the platform can draft a single correlation that replaces them — turning a family of individually unconvincing alerts into one signal that means something.
Candidate multi-stage detections are ranked from your deployed rules, and every candidate is validated before it is shown to you. What did not survive is counted and explained — a short list reads as "three of five failed validation", not as a thin answer.
Every rule is mapped to ATT&CK as it is created, and the coverage matrix ranks what is missing — so the next detection to build is a decision you read off a screen rather than one you argue about. The CVEs on your radar feed the same queue, through the authoring loop.