EXFILSENSE
04 / FRAMEWORK — RANDOM-FOREST DNS_V1 + XGBOOST HOST_V1
measurement: false · combination: false

Simulation under a stated coupling assumption. The score distributions are empirical, taken from the trained models on the TEST split, but the co-occurrence of the two channels is modelled rather than observed: the two branches were trained on unrelated corpora and share no key that would say which DNS record and which host entity belong to one incident. These numbers are not a measurement of a fused detector on paired data.

Read from simulation_manifest.json · split TEST

BranchModel Threshold Rows Positives Checksum
dns random-forest dns_v1 0.4 26 245 7 874 3d456e55dfac6b53…
host xgboost host_v1 0.79492521286 1 548 128 a9dcbb2332e47535…

Variants

The framework pair, and a robustness check that swaps the DNS source. Both are reported so the effect can be seen not to depend on one model choice.

VariantDNSHOST Estimable FN reduction Positive everywhere
framework random-forest dns_v1 xgboost host_v1 36 12.61 % – 80.37 % yes
dns-xgboost-check xgboost dns_v1 xgboost host_v1 36 12.71 % – 80.27 % yes

What was taken from the data and what was assumed. The score distributions are empirical — they are the two trained models' outputs on the TEST split. The co-occurrence of the two channels is not: the branches were trained on unrelated corpora and share no key, so no record on one side can be said to belong to the same incident as a record on the other. That relationship is treated as unknown and swept over the grid below, and at each point a decision threshold is searched.

Manifestation grid
0.5 · 0.65 · 0.8 · 0.9 · 1.0 (5 values)
Coupling grid (ρ)
-0.3 · -0.15 · 0.0 · 0.15 · 0.3 · 0.45 · 0.6 · 0.75 · 0.9 (9 values)
Cells
5 × 9 = 45
Base rate
0.0826873385013
Incidents per draw
200 000
Threshold search points
300
Seed
42

Run

Experiment
fusion-simulation
Model label
DNS x HOST framework
Base version
sim_v2
Simulation time
10.881 s

One cell per point of the grid. Colour is magnitude on a single hue; the number is printed in every cell, so nothing depends on reading a shade. The nine cells the manifest excludes are hatched and marked — they are shown rather than dropped, because which cells could not be estimated is part of the result.

FN reduction 12.61 % → 80.37 % not estimable
False-negative reduction in percent for each manifestation rate and coupling value. Cells marked "not estimable" were excluded.
manifestation \ ρ -0.3-0.150.00.150.30.450.60.750.9
0.5 48.95 46.31 42.92 37.7 33.86 29.63 25.02 18.95 12.61
0.65 64.05 58.91 53.79 49.51 44.39 38.11 32.76 25.34 16.6
0.8 75.25 71.93 65.71 61.66 57.11 48.65 42.88 31.67 21.37
0.9 80.37 76.78 75.23 69.66 65.08 61.47 48.78 40.72 29.76
1.0 not estimable, 6 FN not estimable, 3 FN not estimable, 1 FN not estimable, 2 FN not estimable, 8 FN not estimable, 6 FN not estimable, 1 FN not estimable, 5 FN not estimable, 4 FN

36 of 45 cells estimable. Hover or focus a cell for its detail.

Cells
45
Estimable
36
Minimum FN for a ratio
100
the better single branch produced fewer than 100 false negatives, so a percentage computed from it is dominated by sampling noise

The nine excluded cells and the count that excluded each one. Every value below is under 100, which is the rule quoted above doing its work rather than being asserted.

Manifestation ρ requested Best single branch FN
1.0 -0.3 6
1.0 -0.15 3
1.0 0.0 1
1.0 0.15 2
1.0 0.3 8
1.0 0.45 6
1.0 0.6 1
1.0 0.75 5
1.0 0.9 4
FN reduction
12.61 % – 80.37 % across the 36 estimable cells
Sign of the effect
positive in every estimable cell
Across variants
positive in every variant
Variant lower-bound spread
0.1 percentage points

This is a range over a swept assumption, not an interval around a measurement. No fused detector was evaluated on paired data, so no precision or recall may be quoted for the pair.

The two branches' score distributions and the modelled coupling between them at a single manifestation rate.
The two branches' score distributions and the modelled coupling between them at a single manifestation rate.
False-negative reduction across the whole grid: one line per manifestation rate, swept over the coupling assumption.
False-negative reduction across the whole grid: one line per manifestation rate, swept over the coupling assumption.

In plain language

deterministic

What this model is reacting to. This is a simulation over random-forest dns_v1 + xgboost host_v1. The score distributions are empirical, taken from the TEST split; the co-occurrence of the two channels is modelled, because the branches were trained on unrelated corpora and share no key.

Where to stop. Simulation under a stated coupling assumption. The score distributions are empirical, taken from the trained models on the TEST split, but the co-occurrence of the two channels is modelled rather than observed: the two branches were trained on unrelated corpora and share no key that would say which DNS record and which host entity belong to one incident. These numbers are not a measurement of a fused detector on paired data.

No curated ATT&CK row applies to these features, so no technique is named. The table is referenced, never derived — see why there is no ATT&CK mapping.

Written by
a template no language model was called
Generated
Source artifact
ATT&CK table
123d849fca665e28 draft
Prompt template
2026-08-11.1 current template; not used for this text
Sent to the API
nothing — no request was made. No raw records leave the host.

ANTHROPIC_API_KEY is not set, so the deterministic briefing is shown. This is a supported state, not a failure. Viewing a page never calls the API in either case — generation happens only through python manage.py web build-explanations.

The generated fusion report, served exactly as written.

open in new tab