TID-CMM Threat-Informed Detection Capability Maturity Model
Home › Methodology › ATT&CK Evaluations

Using ATT&CK Evaluations as external evidence

MITRE ATT&CK Evaluations and TID-CMM answer related but different questions. An evaluation result is admissible in some parts of this model and inadmissible in others, and the difference is deliberate.

ATT&CK Evaluations provides independent evidence of how a security product or service performed against defined adversary behaviours in a controlled evaluation scenario. The 2026 methodology scores detection coverage, precision and speed, alongside protection quality and false-positive performance. It also records who produced the result: platform automation, AI, human analysts, a managed service, or a mix.

TID-CMM evaluates something more local: would this organisation actually see the behaviours that matter to it?

An ATT&CK Evaluation result can therefore be used as an external evidence source when assessing a technology’s demonstrated capability against an in-scope technique. It does not establish locally validated coverage.

An evaluation may show that a capability detected an implementation of an ATT&CK technique with high-quality context and low latency. That is useful evidence. TID-CMM still needs to establish whether:

External capability evidence — a technology or service demonstrated that it can detect the behaviour.
Local validated coverage — the organisation demonstrated that it detects the behaviour here.

External evidence can strengthen confidence in a technology selection, explain product capability and guide local validation priorities. It cannot replace local validation.

Where external evidence enters TID-CMM, and where it stopsAn ATT&CK Evaluation result carries a DC tier, a DS tier, a source-of-action modifier and a product and ATT&CK version. It is admissible as evidence in Detection Engineering, informative only in Telemetry and Detection Coverage, and inadmissible in Adversarial Validation, in threat scoping and as proof that telemetry exists locally. Figure 1 — Where external evidence enters TID-CMM, and where it stops An ATT&CK Evaluation result is admissible in one domain, informative in one, and inadmissible everywhere it would inflate a score. ATT&CK Evaluation result per technique · per scenario DC tier · DS tier · modifier product version · ATT&CK version "The capability was demonstrated." TIThreat Intelligence & Adversary Prioritisation scope is chosen by the organisation, not by the evaluation TMThreat Modelling & Attack Path Analysis your paths, your crown jewels DCTelemetry & Detection Coverage informs priority only · never proves telemetry exists here DEDetection Engineering admissible: detection is engineerable, technique-level (DC-3) AVAdversarial Validation & Emulation inadmissible: MITRE's red team did not run in your estate AAAnalytics, Automation & Hunting inadmissible: precision was measured in the range, not here IRIncident Response & Recovery out of scope for detection evidence GVGovernance, Metrics & Continuous Improvement out of scope for detection evidence admissible informs blocked Evidence ceiling Rapid self-assessment: 3.00 Structured: 5.00 · Evidence-based: 5.00 external evidence does not lift it Telemetry evidence the range's data sources are not yours never counts · "a gap in physics" admissible or informative inadmissible — would inflate the score not touched by detection evidence A DC-3 in MITRE's range is evidence of an engineered detection. It is never evidence of a validated one. Validation is local by definition. TID-CMM · © 2022–2026 Reza Adineh · Not affiliated with or endorsed by MITRE. MITRE ATT&CK® is a registered trademark of The MITRE Corporation.
Figure 1 — Where external evidence enters TID-CMM, and where it stops. An ATT&CK Evaluation result is admissible in one domain, informative in one, and inadmissible everywhere it would inflate a score.

Where the evidence lands in the model

TID-CMM has eight domains and five integrity constraints. External evaluation evidence is admissible in some places and inadmissible in others.

It may support DE — Detection Engineering. An evaluation result is evidence that the detection is engineerable and that the product can produce it with the required context. It supports the claim “this detection exists and is technique-level”. It does not support the claim “this detection has fired correctly here”.

It may inform DC — Telemetry & Detection Coverage. MITRE’s per-technique results show which behaviours a product can surface given the telemetry the range provided. That is useful for ranking which telemetry to enable next. It is not evidence that the telemetry exists in your estate.

It is never evidence in AV — Adversarial Validation. Validation is the price of the upper levels, and MITRE’s red team did not run in your environment. A DC-3 in MITRE’s range is evidence of an engineered detection, never of a validated one.

It never counts as telemetry evidence. The range’s data sources are not yours. If Sysmon-class or EDR telemetry is absent locally, no evaluation result changes that. That is a gap in physics.

Protection results are advisory in the same way. A behaviour a product blocked in the range is evidence about the product, not about your estate; the prevention dimension of AV.8 is filled only by a local emulation. See prevention, protection and the TIR-CMM boundary.

It does not lift the evidence ceiling. A rapid self-assessment with a folder of evaluation results attached is still capped at 3.00 by C3. Ceilings are set by the assessment depth and the quality of local evidence. External evidence is context, not depth.

Reading MITRE’s tiers in TID-CMM terms

MITRE scores each behaviour on a Detection Coverage tier. Those tiers map closely onto TID-CMM’s telemetry assurance bands — with one important difference at the top.

MITRE DC tierWhat it means in the range Closest TID-CMM stateWhat it can and cannot support
DC-0 Not observedNo alert, no telemetryBlindEvidence that the product did not see it there. Not proof it cannot be seen here with different telemetry.
DC-1 ObservedTelemetry exists, no alertTelemetry assured, detection absentThe distinction TID-CMM is built on. Evidence that observation is possible; no detection to credit.
DC-2 Correlated (generic)Tactic-level alert, needs enrichmentPartial detectionSupports a generic or partial DE claim. Does not support technique-level coverage.
DC-3 ActionableTechnique-level alert with WHO, WHAT, WHEN, WHERE, HOW, SEVERITYEngineered, technique-levelSupports an engineered DE claim. Never maps to validated. Validation is local.
N/A Not assessedNot in scope or not conducted—No evidence. Not negative evidence.

MITRE’s DC-3 alert content — WHO, WHAT, WHEN, WHERE, HOW, SEVERITY — is also a good working definition of “actionable” for the DE domain. An alert that cannot answer those six questions is not technique-level, whoever produced it.

Four things the evaluation cannot tell you

Scenario scope. MITRE tested one procedure of each technique, in one scenario, with weights set for that adversary’s objective. Your Tier A set will overlap that scenario only partially. A DC-3 on one procedure of T1003.001 is not coverage of every procedure of T1003.001. A technique absent from the evaluation, or marked N/A, is no evidence — not negative evidence, and not positive evidence either.

Who produced it. Every 2026 score carries a Source-of-Action modifier: (P) platform, (P/AI) platform with AI, (S/H) service with human analysts, (S/AI) AI-driven service, or Mixed. If the result was produced (S/H) — the vendor’s own analysts enriching and escalating — and the organisation is buying the platform without the service, the (S/H) result is not evidence for the deployment being assessed. Only evidence produced under the operating model you are actually deploying counts. Record the modifier. Always.

Precision. MITRE’s Detection Precision is measured against the range’s benign activity and includes a case-consolidation penalty. It says nothing about signal-to-noise in your estate. Local precision is measured locally, in the AA domain.

Freshness. Evaluations are per round. The product version tested is not the version deployed eighteen months later, and the ATT&CK version the round used may not be v19.2 — technique IDs split and merge between versions. External evidence expires like any other evidence. It carries a round, a product version and an ATT&CK version, and it leaves the validation window when any of those no longer matches what is deployed.

A practical use

For every Tier A technique derived from the organisation’s platforms, adversaries and attack paths, the assessment records relevant ATT&CK Evaluations evidence alongside its local evidence.

The per-technique evidence chain, and what a mismatch revealsFor each Tier A technique the chain runs from the relevant behaviour, through required telemetry, local detection and local validation. External evaluation evidence sits alongside it and is advisory; local evidence is authoritative. Where the two disagree, the disagreement is the finding. Figure 2 — The per-technique evidence chain, and what a mismatch reveals For every Tier A technique, external evidence sits beside local evidence. When they disagree, the disagreement is the finding. Relevant behaviour Tier A · from your attack paths External evidence ATT&CK Evaluation result Required telemetry exists on scope · band Local detection deployed · configured · engineered Local validation reproduced here · in window advisory authoritative External record what MITRE demonstrated, and under what conditions Round / scenarioEnterprise 2026 · scenario 2 Product / version testedVendor X EDR 7.4 ATT&CK version of the roundv19 DC tier · DS tierDC-3 · Real-time (<15 min) Source-of-Action modifier(P)must match the operating model you deploy N/A flagno expires when product version, ATT&CK version or round no longer matches what is deployed Local record what the organisation demonstrated, here Telemetry bandassured · partial · weak · blind Deployed on scopeyes · partial · no Engineeredyes · no Validatedyes · no Validation date · window2026-07-12 · 90 days Local precision (AA)measured here, never imported this is the record that sets the score Strong external · failed local Not automatically a product problem. Look for: missing telemetry · incomplete deployment configuration drift · integration failure pipeline latency · handoff failure an untested local assumption this is the distinction TID-CMM exists to expose Weak or absent external · passing local A detection the organisation engineered itself, which the vendor never demonstrated. Local evidence wins outright. Absence from the evaluation is no evidence — not negative, not positive. TID-CMM · © 2022–2026 Reza Adineh · Not affiliated with or endorsed by MITRE. MITRE ATT&CK® is a registered trademark of The MITRE Corporation.
Figure 2 — The per-technique evidence chain, and what a mismatch reveals. For every Tier A technique, external evidence sits beside local evidence. When they disagree, the disagreement is the finding.

The external record for a technique holds:

FieldExample
Evaluation round and scenarioEnterprise 2026, scenario 2
Product and version testedVendor X EDR 7.4
ATT&CK version used by the roundv19
DC tierDC-3
DS tierReal-time (<15 min)
Source-of-Action modifier(P)
N/A flagno

Next to it, the local record holds the fields TID-CMM already uses: telemetry band, deployed on scope (yes / partial / no), engineered (yes / no), validated (yes / no), validation date, and the validation window.

That creates a useful comparison for each technique: relevant behaviour → external capability evidence → required telemetry → local detection → local validation.

Where strong external evidence exists but local validation fails, the gap is no longer automatically a product problem. It may instead expose:

This is exactly the distinction TID-CMM is intended to expose.

The mirror case matters too. Where external evidence is weak or absent but local validation passes — a detection the organisation engineered itself, which the vendor never demonstrated — the local result wins outright. External evidence is advisory. Local evidence is authoritative.

The rule

ATT&CK Evaluations should be consumed as evidence, not converted into a TID-CMM score.

TID-CMM remains organisation-specific and threat-scoped. Its validated coverage reflects what has been demonstrated inside the assessed environment, not what a technology demonstrated somewhere else.

How TID-CMM consumes ATT&CK · What counts as evidence · The five constraints · Prevention and the TIR-CMM boundary

MITRE ATT&CK® is a registered trademark of The MITRE Corporation. TID-CMM is an independent project and is not affiliated with or endorsed by MITRE.

Diagram

100%