Platform · 01 Ingest

Any source. Any environment.
Zero parser projects.

If it emits logs, Qvasir can consume it — cloud services, on-prem servers, network gear, operating systems, applications, containers. Deploy the platform wherever your data lives: public cloud, private datacenter, or air-gapped, with ingest, detection and search all running offline.

Onboarding a new source is a guided workflow, not an integration project.

A new log source is normally a project: someone reads the vendor's documentation, writes a parser, guesses at a field mapping, and finds out months later which parts were wrong. Here it is five steps, and you only perform two of them.

01 · Collect

Point a collector at it

Lightweight open-source collection (Vector) ships events over a Kafka-compatible backbone — buffered, reliable and replayable.

02 · Sample

Find every shape in the stream

Diversity-aware sampling looks for the distinct event shapes a topic carries, rather than the first N events — so a rare format is not missed because it is rare.

03 · Identify

What is this, and who made it?

The vendor and product are identified from the events themselves, and each cluster of events is described in terms of what it actually records.

04 · Classify

OCSF class and entities

Each cluster is classified into the OCSF taxonomy, and the fields carrying users, hosts and addresses are proposed as entities with the role each one plays.

05 · Approve

You decide what is right

Every proposal is presented for review, correctable, and nothing is written to the catalogue until you accept it.

Steps 03 and 04 are four separate model passes — vendor and product identification, cluster enrichment, OCSF classification, and entity proposal. That is the part that used to be a parser project, and it is the reason a source can go from first event to usable schema in an afternoon. Your job is step 05: reading what it concluded and correcting it where it is wrong. Detections then bind to your real field names, not to a normalised copy — the OCSF class is how coverage and correlation talk about the source, not what your rules match on.

Source definition for a Windows System log: entity keyability measured over 1,534 events, 66 of 66 fields observed, each with data type, entity role and live samples
Every source's discovered schema, measured: which entities the events can key on and how often, every observed field with its type, role and live samples.

One dependency, stated plainly. Because onboarding is AI-assisted end to end, it needs a reachable model provider — today that means one of the supported cloud APIs. Ingest, detection, correlation and search all run offline, but a fully disconnected deployment cannot currently add a new source. We would rather you knew that before the pilot than during it.

Roadmap Support for self-hosted open-weight models is on the roadmap, and it is the change that closes this gap — pointing the platform at a model running inside your own network, so onboarding works with nothing leaving it. See the roadmap →

What actually gets installed

Three ways in, and one we do not pretend to have yet.

Vector, at the source

Lightweight open-source agents run where the logs are — tailing files, receiving syslog, reading the Windows event log — and ship to the streaming backbone Qvasir deploys with. Vector buffers to disk, so a network outage delays events rather than dropping them.

Cloud sources, pulled directly

Where there is no host to install anything on, the source is read from its own API instead. Microsoft Entra ID and Google Workspace are collected this way today — no agent, no forwarder, and the same onboarding workflow as everything else.

Or write to the topic yourself

The backbone is Kafka-compatible, so anything already speaking that protocol can produce to it — a shipper you already run, an existing pipeline, or a service emitting directly. Adopting our collector is not a condition of using the platform.

Your own Kafka estate Roadmap

Reading from brokers you already operate — your authentication, your topics, your ACLs — is not supported today. Qvasir consumes from the backbone it ships with. If you run Kafka already and want it read in place, tell us: it is a priority set by who asks for it.

Data sources in a demonstration estate: seventeen sources — Google Workspace, Home Assistant, Linux auditd, Entra ID, Windows DNS, Defender Firewall, PowerShell, Security and System logs, Proxmox — each showing event types, field count, rules bound to it, engine state and health
Seventeen sources in a demonstration estate — cloud, identity, Windows, Linux, network, virtualisation — each with its field count, the rules bound to it, and its health.

Your existing tools

Your EDR is a source, not a competitor.

Replacing the SIEM layer does not mean replacing the tools that feed it. The security products you already run — EDR, firewall, email security, cloud posture, identity protection — onboard through the same guided workflow as any other source, and their alerts are treated as what they are: high-quality signals from a narrow vantage point.

Alerts become signals

A vendor alert enters the same pipeline as raw telemetry — classified into OCSF, resolved to the entities behind it, and scored alongside everything else those entities did.

Correlated across vendor boundaries

Your EDR sees the endpoint. Your identity provider sees the login. Your firewall sees the egress. None of them sees the incident. Qvasir correlates across all three — and the raw telemetry underneath them.

Their strength, kept

A good EDR detection is worth having. It arrives with its verdict and context intact and becomes evidence inside an investigation, instead of another console somebody has to remember to check.

This is one half of the cost argument, and the halves are easy to confuse. What gets replaced is the expensive middle — the SIEM licence, the SOAR project, the MDR retainer that stores, automates and reads. What stays is everything that produces signal. See the side-by-side against SIEM and MDR models, or what that distinction means to a budget holder.

Your telemetry changes. Qvasir notices.

Schema drift tracking

When a vendor update adds, renames, or drops fields, Qvasir detects the drift, surfaces it, and guides the correction — before detections silently go blind.

Measured field statistics

The platform continuously measures which fields are actually present in your data and how they're populated — so detection authoring is grounded in reality, not vendor documentation.

Ingestion health, honestly reported

Per-source throughput, freshness, and health are observable at a glance. When something is unavailable, Qvasir says so — it never fabricates a green light.

Ingestion and storage observability: 14 healthy and 3 learning sources, 81.7 million rows on 15.4 GiB at 12.7x compression, events per second per topic over 24 hours, and producer-side counters for buffer shedding, socket loss, under-replicated partitions and Kafka to ClickHouse divergence
Ingestion as reported, not assumed: events per topic over the last day, and the loss counters at every hop between the collector and the store.