Platform · 04 Improve

The alert you close at 03:00
is one you stop seeing.

Detection quality decays: sources change, environments drift, and the rule that was precise in March is noise by August. Qvasir feeds real outcomes back into the rules that produced them, so tuning stops being a quarterly project and becomes a property of the system.

Built-in improvement loops, not a dashboard of regrets.

  • Noisy rules get proposed refinements. Match volumes and false-positive patterns identify the rules costing you the most attention, with a concrete suggested change rather than a number on a chart.
  • Silent rules get checked against reality. A rule that never fires is either excellent or broken, and those look identical on a dashboard. Qvasir distinguishes them against the schema and the events the source is actually carrying.
  • Your verdict is an input, not paperwork. Analyst feedback and the agents' own false-positive findings both open concrete tuning suggestions against the offending rule.
  • Suggestions arrive with their reasoning. Every proposed change carries the evidence behind it, the rationale, and a confidence — so approving one is a judgement, not a leap.
Tuning queue with 26 open suggestions from the engine's false-positive findings and match-rate guardrails — each naming the detection rule, the proposed whitelist value or logic change, and its measured effect such as stops 433 of 433 or no effect on 1 signal
The tuning queue — every suggestion names its rule, its proposed change, and what it would have stopped, measured before you accept it.

Calibration

By the time you open the rule, the numbers are already there.

Tuning normally starts with a blank tab and a decision to go and measure something. Here the measuring has already happened: save a generated rule and its calibration against your own history is dispatched without being asked, so what you arrive at is a false-positive surface rather than an empty panel and an intention.

Exclusions are proposed, with the vote shown

Candidate values to exclude are not one model's opinion. Each is judged three times and aggregated, and you see the level of agreement alongside the proposal — so a unanimous call and a narrow one do not look identical on the screen.

It cannot exclude the thing that fired

A proposal that would whitelist the rule's own trigger value is rejected in code, whatever the model concluded. That is the one exclusion that would silently neuter the rule, so it is not left to judgement.

Too many values is an answer, not a list

Past a threshold of distinct values, the platform advises against excluding at all — deterministically. A field with hundreds of values is telling you the rule is wrong, not that you need a longer allowlist.

The guardrail that matters

The AI can make this more cautious. It cannot make it look safer.

This is the part worth reading carefully, because it is a design decision rather than a feature. Wherever a model touches a live detection, what it is permitted to do is constrained by code — not by an instruction in a prompt that a cleverly-worded log line might talk it out of.

A one-way ratchet on trust

The AI reviewing your calibration results can demote a value to "review carefully" or attach a warning. It structurally cannot approve a value, raise a confidence score, or change any number — the merge is an add-only operation. And if a merge ever ended up relaxing the deterministic verdict, the entire AI contribution is discarded rather than trusted. It can only ever narrow.

Scope-locked to the part it was asked about

When it refines a rule from a tuning suggestion, it may change the detection logic and the tunables. Anything it returns outside that scope is reverted before you ever see it — so "fix the false positives" cannot quietly become a rewritten severity, a changed source binding, or a new set of ATT&CK tags.

It is allowed to disagree with you

Hand it a tuning suggestion and it can come back rejecting it, per suggestion, with a reason. An assistant that implements every instruction including the bad ones is not much of a safeguard.

It never sees the raw events

Calibration data contains values an attacker can influence — a hostname, a filename, a command line. Those are fenced inside a delimited block that is regenerated per run, and the model reasons over aggregates rather than the raw stream. Prompt injection through your own telemetry is a real attack on a platform like this one.

The pattern underneath all four is the same: the AI proposes, deterministic code decides what a proposal is even allowed to be, and a person applies the ones that reach a live system. Every applied change lands in the rule's own version history.

When tuning is not the answer

Sometimes the fix is fewer rules, not quieter ones.

Noisy rules become one correlation

When a group of rules is noisy for the same underlying reason, the platform can draft a single correlation that replaces them — turning a family of individually unconvincing alerts into one signal that means something.

It proposes chains from what you already run

Candidate multi-stage detections are ranked from your deployed rules, and every candidate is validated before it is shown to you. What did not survive is counted and explained — a short list reads as "three of five failed validation", not as a thin answer.

The backlog writes itself

Every rule is mapped to ATT&CK as it is created, and the coverage matrix ranks what is missing — so the next detection to build is a decision you read off a screen rather than one you argue about. The CVEs on your radar feed the same queue, through the authoring loop.