Using ATT&CK Evaluations as external evidence
MITRE ATT&CK Evaluations and TID-CMM answer related but different questions. An evaluation result is admissible in some parts of this model and inadmissible in others, and the difference is deliberate.
ATT&CK Evaluations provides independent evidence of how a security product or service performed against defined adversary behaviours in a controlled evaluation scenario. The 2026 methodology scores detection coverage, precision and speed, alongside protection quality and false-positive performance. It also records who produced the result: platform automation, AI, human analysts, a managed service, or a mix.
TID-CMM evaluates something more local: would this organisation actually see the behaviours that matter to it?
An ATT&CK Evaluation result can therefore be used as an external evidence source when assessing a technology’s demonstrated capability against an in-scope technique. It does not establish locally validated coverage.
An evaluation may show that a capability detected an implementation of an ATT&CK technique with high-quality context and low latency. That is useful evidence. TID-CMM still needs to establish whether:
- the technique is relevant to the organisation’s threat model and attack paths;
- the required telemetry exists across the relevant estate;
- the capability is actually deployed on that scope;
- the local configuration preserves the evaluated capability;
- the detection reaches the correct operational destination;
- the behaviour has been reproduced or otherwise validated locally;
- the evidence remains inside its required validation window.
External capability evidence — a technology
or service demonstrated that it can detect the behaviour.
Local validated coverage — the organisation demonstrated that it
detects the behaviour here.
External evidence can strengthen confidence in a technology selection, explain product capability and guide local validation priorities. It cannot replace local validation.
Where the evidence lands in the model
TID-CMM has eight domains and five integrity constraints. External evaluation evidence is admissible in some places and inadmissible in others.
It may support DE — Detection Engineering. An evaluation result is evidence that the detection is engineerable and that the product can produce it with the required context. It supports the claim “this detection exists and is technique-level”. It does not support the claim “this detection has fired correctly here”.
It may inform DC — Telemetry & Detection Coverage. MITRE’s per-technique results show which behaviours a product can surface given the telemetry the range provided. That is useful for ranking which telemetry to enable next. It is not evidence that the telemetry exists in your estate.
It is never evidence in AV — Adversarial Validation. Validation is the price of the upper levels, and MITRE’s red team did not run in your environment. A DC-3 in MITRE’s range is evidence of an engineered detection, never of a validated one.
It never counts as telemetry evidence. The range’s data sources are not yours. If Sysmon-class or EDR telemetry is absent locally, no evaluation result changes that. That is a gap in physics.
Protection results are advisory in the same way. A behaviour a product
blocked in the range is evidence about the product, not about your estate; the prevention
dimension of AV.8 is filled only by a local emulation. See
prevention, protection and the
TIR-CMM boundary.
It does not lift the evidence ceiling. A rapid self-assessment with a folder of evaluation results attached is still capped at 3.00 by C3. Ceilings are set by the assessment depth and the quality of local evidence. External evidence is context, not depth.
Reading MITRE’s tiers in TID-CMM terms
MITRE scores each behaviour on a Detection Coverage tier. Those tiers map closely onto TID-CMM’s telemetry assurance bands — with one important difference at the top.
| MITRE DC tier | What it means in the range | Closest TID-CMM state | What it can and cannot support |
|---|---|---|---|
| DC-0 Not observed | No alert, no telemetry | Blind | Evidence that the product did not see it there. Not proof it cannot be seen here with different telemetry. |
| DC-1 Observed | Telemetry exists, no alert | Telemetry assured, detection absent | The distinction TID-CMM is built on. Evidence that observation is possible; no detection to credit. |
| DC-2 Correlated (generic) | Tactic-level alert, needs enrichment | Partial detection | Supports a generic or partial DE claim. Does not support technique-level coverage. |
| DC-3 Actionable | Technique-level alert with WHO, WHAT, WHEN, WHERE, HOW, SEVERITY | Engineered, technique-level | Supports an engineered DE claim. Never maps to validated. Validation is local. |
| N/A Not assessed | Not in scope or not conducted | — | No evidence. Not negative evidence. |
MITRE’s DC-3 alert content — WHO, WHAT, WHEN, WHERE, HOW, SEVERITY — is also a good working definition of “actionable” for the DE domain. An alert that cannot answer those six questions is not technique-level, whoever produced it.
Four things the evaluation cannot tell you
Scenario scope. MITRE tested one procedure of each technique, in one scenario, with weights set for that adversary’s objective. Your Tier A set will overlap that scenario only partially. A DC-3 on one procedure of T1003.001 is not coverage of every procedure of T1003.001. A technique absent from the evaluation, or marked N/A, is no evidence — not negative evidence, and not positive evidence either.
Who produced it. Every 2026 score carries a Source-of-Action modifier: (P) platform, (P/AI) platform with AI, (S/H) service with human analysts, (S/AI) AI-driven service, or Mixed. If the result was produced (S/H) — the vendor’s own analysts enriching and escalating — and the organisation is buying the platform without the service, the (S/H) result is not evidence for the deployment being assessed. Only evidence produced under the operating model you are actually deploying counts. Record the modifier. Always.
Precision. MITRE’s Detection Precision is measured against the range’s benign activity and includes a case-consolidation penalty. It says nothing about signal-to-noise in your estate. Local precision is measured locally, in the AA domain.
Freshness. Evaluations are per round. The product version tested is not the version deployed eighteen months later, and the ATT&CK version the round used may not be v19.2 — technique IDs split and merge between versions. External evidence expires like any other evidence. It carries a round, a product version and an ATT&CK version, and it leaves the validation window when any of those no longer matches what is deployed.
A practical use
For every Tier A technique derived from the organisation’s platforms, adversaries and attack paths, the assessment records relevant ATT&CK Evaluations evidence alongside its local evidence.
The external record for a technique holds:
| Field | Example |
|---|---|
| Evaluation round and scenario | Enterprise 2026, scenario 2 |
| Product and version tested | Vendor X EDR 7.4 |
| ATT&CK version used by the round | v19 |
| DC tier | DC-3 |
| DS tier | Real-time (<15 min) |
| Source-of-Action modifier | (P) |
| N/A flag | no |
Next to it, the local record holds the fields TID-CMM already uses: telemetry band, deployed on scope (yes / partial / no), engineered (yes / no), validated (yes / no), validation date, and the validation window.
That creates a useful comparison for each technique: relevant behaviour → external capability evidence → required telemetry → local detection → local validation.
Where strong external evidence exists but local validation fails, the gap is no longer automatically a product problem. It may instead expose:
- missing telemetry;
- incomplete deployment;
- configuration drift;
- integration failure;
- detection pipeline latency;
- operational handoff failure;
- or an untested local assumption.
This is exactly the distinction TID-CMM is intended to expose.
The mirror case matters too. Where external evidence is weak or absent but local validation passes — a detection the organisation engineered itself, which the vendor never demonstrated — the local result wins outright. External evidence is advisory. Local evidence is authoritative.
The rule
ATT&CK Evaluations should be consumed as evidence, not converted into a TID-CMM score.
TID-CMM remains organisation-specific and threat-scoped. Its validated coverage reflects what has been demonstrated inside the assessed environment, not what a technology demonstrated somewhere else.
How TID-CMM consumes ATT&CK · What counts as evidence · The five constraints · Prevention and the TIR-CMM boundary
MITRE ATT&CK® is a registered trademark of The MITRE Corporation. TID-CMM is an independent project and is not affiliated with or endorsed by MITRE.