4. Structured Analysis of Adversary Behavior

Reasoning Under Uncertainty

Recognize assumptions and cognitive bias, use estimative language consistently, and make confidence judgments that reflect evidence quality and analytic agreement.

In this lesson, you will learn to:

  • Use calibrated language, explicit assumptions, and confidence judgments to explain what is known, what is uncertain, and what evidence would change the assessment.

Reasoning Under Uncertainty

Separate facts, assumptions, and judgments

Intelligence analysis turns incomplete and sometimes contradictory evidence into judgments that support decisions. The analyst cannot remove uncertainty, but can make the reasoning visible: what was observed, what was assumed, what was inferred, which alternatives remain, and what evidence could change the conclusion.

Start by separating four components:

Component Meaning Northbridge example
Fact or sourced claim A statement that can be checked against a source or observation Three finance workstations launched the same command interpreter within ten minutes of opening archive attachments.
Assumption A condition accepted as true for the analysis but not fully established Endpoint clocks are sufficiently synchronized to compare the event sequence.
Analytic judgment A conclusion derived from evidence and reasoning The events were likely related rather than independent user actions.
Implication What the judgment could mean for a consumer or decision Related credential activity may justify widening the investigation beyond the three workstations.

A sourced statement is not necessarily true merely because it is written as a fact. Record its provenance and limitations. Conversely, a judgment is not defective because it cannot be observed directly; inference is the analyst’s work. The requirement is to label and support it.

Build an evidence-to-judgment chain

A defensible judgment should allow another analyst to reconstruct the path:

  1. State the proposition. What precisely are you trying to determine?
  2. List relevant evidence. Include supporting, contradictory, and missing information.
  3. Expose assumptions. Which unverified conditions connect the evidence to the proposition?
  4. Generate alternatives. What other explanations could plausibly produce the observations?
  5. Compare explanations. Which evidence is expected, inconsistent, or non-discriminating under each explanation?
  6. Form the judgment. Which explanation currently fits best, and how likely is it?
  7. Assess confidence. How strong is the evidentiary and reasoning foundation?
  8. State implications and change conditions. Why does the judgment matter, and what new evidence would revise it?

Suppose Northbridge observes an unusual process sequence on three finance workstations. The proposition is: The systems are affected by one coordinated intrusion attempt.

Supporting evidence includes similar email attachments, closely timed execution, the same process ancestry, and connections to the same newly observed domain. Contradictory or limiting evidence includes the absence of confirmed payload execution on one host and a legitimate software package that can produce part of the sequence. Missing evidence includes memory captures and complete identity telemetry.

The judgment should not leap from “same domain” to “coordinated intrusion.” It should explain why the combined sequence is more consistent with coordination than with benign software, independent user behavior, or a logging artifact.

Make assumptions explicit and testable

Assumptions fill gaps that cannot immediately be resolved. Hidden assumptions are dangerous because they can carry most of a conclusion’s weight without being examined.

Common CTI assumptions include:

  • telemetry is sufficiently complete to reveal relevant activity;
  • timestamps from different systems can be compared;
  • an account’s recorded identity reflects the person or process controlling it;
  • external reporting describes the same tool, behavior, or actor definition used internally;
  • an observed infrastructure relationship remained valid during the assessed period;
  • affected organizations reported incidents consistently;
  • the consumer’s asset inventory and criticality labels are current;
  • an adversary’s past behavior is informative about future behavior.

For each key assumption, record:

  • why it is necessary;
  • which evidence supports it;
  • how sensitive the judgment is to it;
  • how it could be tested;
  • who can validate it;
  • when it should be reviewed.

A load-bearing assumption is one whose failure would materially change the judgment. Prioritize collection against load-bearing assumptions rather than trying to verify every minor detail.

For example, if the coordinated-intrusion judgment depends on event order, clock synchronization is load-bearing. If clock error could reverse the sequence, the analyst should verify offsets or lower confidence.

Generate alternatives before settling on a story

People naturally seek coherence. Once a plausible explanation appears, analysts may interpret later evidence as confirmation. Counter this tendency by generating alternatives early.

For the Northbridge activity, alternatives might include:

  • coordinated malicious activity delivered through similar messages;
  • an authorized software deployment that produced unfamiliar behavior;
  • independent user actions coinciding in time;
  • a security test or administrative task not communicated to monitoring teams;
  • duplicated or misattributed telemetry;
  • a mixture of benign and malicious events rather than one common cause.

Alternatives should be plausible enough to test, not an exhaustive list of imaginable possibilities. Avoid a false binary such as “named threat group or benign.” Multiple malicious, benign, and mixed explanations may fit.

Ask of each alternative:

  • What would we expect to observe if it were true?
  • Which current evidence is difficult to explain?
  • Which evidence is compatible with several alternatives and therefore weakly discriminating?
  • What obtainable evidence would best distinguish it from the leading explanation?
Recognize cognitive traps

Cognitive bias cannot be eliminated through willpower, but analytic process can reduce its influence.

Trap How it appears Countermeasure
Anchoring The first report, actor label, or hypothesis controls later interpretation. Delay commitment; record an initial evidence inventory and alternatives.
Confirmation bias Analysts seek support for the favored explanation and explain away contradictions. Assign someone to identify disconfirming evidence and specify change conditions.
Availability bias A recent or memorable incident seems more likely than base conditions justify. Compare against internal prevalence, baselines, and relevant reference classes.
Premature closure Collection stops after one plausible explanation emerges. Use an analysis gate that requires alternatives and remaining gaps.
Mirror imaging The adversary is assumed to share the analyst’s incentives, constraints, or risk tolerance. Separate observed behavior from inferred intent and consider different objectives.
Attribution bias Activity is explained through actor identity while environmental or technical causes receive less attention. Analyze behavior and opportunity before attaching identity labels.
Groupthink Team agreement is mistaken for evidentiary strength. Use independent review, structured dissent, or anonymous initial judgments.
Base-rate neglect A suspicious feature outweighs how commonly it occurs in benign activity. Establish prevalence and compare conditional likelihoods.
Outcome bias A judgment is considered good or bad only because of what happened afterward. Evaluate whether the process and evidence were reasonable at the time.

Bias language should improve method, not become a way to dismiss colleagues. Saying “you are biased” rarely helps. Ask instead which alternative was tested, which evidence would disconfirm the judgment, or whether a relevant baseline exists.

Use diagnostic evidence

Evidence is most useful when it distinguishes among explanations. An observation can strongly confirm that something happened while contributing little to which explanation is best.

A newly registered domain may be compatible with malicious delivery, testing, marketing, or ordinary service deployment. It increases suspicion only in combination with context. A rare process sequence following a targeted attachment may discriminate more strongly if benign software rarely produces it.

Build a compact comparison:

Evidence Coordinated intrusion Authorized deployment Logging artifact
Similar attachments reached finance users Expected Difficult to explain Could be real but unrelated to process records
Same unusual process ancestry Expected Possible if deployment package uses it Possible if parsing duplicated fields
Matching destination shortly after execution Expected Possible but requires approved service Difficult to explain if raw records confirm distinct connections
Change ticket or software owner confirmation Unexpected unless coincidental Strongly expected Neutral
Raw event identifiers from separate hosts Expected Expected Contradicts simple duplication

The table does not calculate truth mechanically. It identifies which evidence warrants attention and where additional collection will have the greatest value.

Distinguish absence of evidence from evidence of absence

A missing observation supports absence only when the collection system was capable of detecting the activity under the relevant conditions.

Before writing “no credential access occurred,” ask:

  • Was suitable identity telemetry enabled?
  • Did it cover the relevant accounts, systems, and time period?
  • Was the data retained and successfully queried?
  • Would the suspected technique produce a visible event?
  • Could the activity avoid or tamper with the source?

If coverage is incomplete, write: “We found no credential-access evidence in the available identity and endpoint records; visibility is incomplete for legacy systems, so credential access cannot be ruled out.”

This is not evasive language. It precisely describes what the evidence can support.

Record what would change the assessment

A useful judgment is revisable. State indicators that would strengthen, weaken, or overturn it.

For Northbridge:

  • Would strengthen: decoded payload similarity across hosts, a shared identity sequence, or independent evidence connecting the destination to the same delivery behavior.
  • Would weaken: confirmation of an authorized deployment matching the complete sequence or evidence that events were generated during a sanctioned exercise.
  • Would overturn: validated records proving the apparent events resulted from duplicated telemetry rather than activity on the hosts.

Change conditions guide collection and make updates intelligible to consumers. They also reduce the tendency to defend an earlier judgment after the evidence changes.

Worked analytic note

A concise internal note could read:

Judgment: The three finance workstations were likely affected by a coordinated intrusion attempt rather than unrelated user activity.

Evidence: Each user received a similar archive attachment; the hosts produced the same unusual process sequence and contacted the same destination within a narrow period. Raw records confirm the events originated from distinct hosts.

Assumptions: Host clocks are comparable within two minutes, and the process records have not been altered. Clock validation supports but does not completely establish the first assumption.

Alternatives: An authorized deployment remains possible but is less consistent with the targeted messages and lack of a change record. A telemetry artifact is unlikely because raw identifiers differ across hosts.

Gaps: Payload execution and subsequent credential activity are not confirmed. Identity coverage is incomplete on one legacy service.

Change conditions: Owner confirmation of a matching approved deployment would lower likelihood substantially; matching payload or identity evidence would increase it.

The note separates evidence from inference without burying the reader in process.

Key takeaways
  • Facts, assumptions, judgments, and implications play different roles and should be distinguishable.
  • A defensible assessment shows the chain from proposition through evidence and alternatives to judgment.
  • Identify and test load-bearing assumptions first.
  • Generate plausible alternatives before evidence is organized into one compelling story.
  • Seek diagnostic and disconfirming evidence, not only additional support.
  • Absence becomes evidence only when collection was capable of revealing the activity.
  • State what would change the assessment so analysis remains responsive to new evidence.

Analyst habit: Before finalizing a judgment, ask: Which part of my reasoning would a skeptical, well-informed reviewer challenge first? Address that point explicitly.

Express likelihood and confidence with discipline

Analysts communicate two different ideas when describing uncertainty:

  • Likelihood expresses how probable the analyst judges an event, explanation, or outcome to be.
  • Confidence expresses how strongly the evidence and reasoning support that likelihood judgment.

Do not use the terms interchangeably. An outcome can be likely with low confidence when the available evidence points in one direction but is sparse or unreliable. An outcome can be unlikely with high confidence when strong evidence consistently weighs against it.

Judgment Meaning
Likely, high confidence Strong evidence and sound reasoning support a probability above the chosen threshold.
Likely, low confidence The leading interpretation is more probable than not, but important evidence gaps or weaknesses remain.
Unlikely, high confidence Strong evidence indicates the proposition probably is not true.
Unlikely, low confidence Available evidence weighs against the proposition, but visibility is too limited for a firm conclusion.

This distinction prevents a common ambiguity: “We have low confidence” does not tell the consumer whether the event itself is judged unlikely or whether the analyst is uncertain about a judgment that could still be likely.

Establish a shared likelihood vocabulary

Words such as possible, likely, and almost certain mean different things to different readers. Organizations should define a small estimative vocabulary and use it consistently.

One illustrative scale is:

Term Approximate probability range
Almost no chance 1–5%
Very unlikely 5–20%
Unlikely 20–40%
Roughly even chance 40–60%
Likely 60–80%
Very likely 80–95%
Almost certain 95–99%

The exact boundaries are organizational choices, not universal truths. Publish the definitions, train analysts and consumers to use them, and avoid mixing incompatible scales within one product.

Probability ranges are calibration aids. They do not imply that analysts calculated a precise statistical result. Most CTI judgments combine incomplete observations, source evaluation, comparison of alternatives, and expert reasoning. Use numbers only when the underlying data and method justify them.

Avoid ambiguous modifiers such as may, could, potentially, and cannot rule out when making the main likelihood judgment. Nearly any event could occur. These phrases often describe possibility without helping the consumer compare outcomes.

Prefer:

We assess the activity is likely part of a coordinated intrusion attempt.

Over:

The activity could potentially be related to an intrusion.

The second statement avoids a useful judgment unless mere possibility is itself decision-relevant.

Base confidence on evidence and reasoning

Confidence should reflect the quality of the analytic foundation, not the analyst’s personality or rhetorical certainty. Evaluate at least four dimensions:

  1. Evidence quality: Are important sources reliable, claims credible, and observations relevant and timely?
  2. Evidence sufficiency: Is there enough information to address the proposition, including necessary coverage and context?
  3. Corroboration and consistency: Do independent evidence streams agree, and are contradictions understood?
  4. Reasoning quality: Were assumptions, alternatives, biases, and diagnostic evidence examined?

An illustrative confidence framework is:

Confidence Typical conditions
High Strong, relevant evidence from suitable and substantially independent sources; limited important contradictions; key assumptions tested; alternatives examined; reasoning robust to plausible changes.
Moderate Credible and relevant evidence supports the judgment, but some gaps, dependencies, or unresolved contradictions could change its strength or scope.
Low Evidence is fragmented, weak, indirect, poorly corroborated, or dependent on major assumptions; alternatives remain difficult to distinguish.

Confidence should accompany the judgment and its rationale:

We assess the three workstation events were likely coordinated, with moderate confidence. Similar delivery, process, and network sequences support a common cause, but payload execution is unconfirmed and identity visibility is incomplete.

The explanation matters more than the label. It tells the consumer why confidence is not higher and directs further collection.

Do not average unlike weaknesses

A single confidence score can hide different problems. Two judgments may both receive “moderate confidence” for very different reasons:

  • one has abundant evidence from sources with uncertain independence;
  • another has two highly reliable observations but limited environmental coverage;
  • a third has good evidence but depends on a load-bearing assumption about timing.

Record the reason for the confidence level. Where the distinction matters, identify the limiting dimension directly: source access, coverage, timeliness, contradiction, assumption sensitivity, or analytic disagreement.

Confidence is also proposition-specific. The team may have high confidence that exploitation occurred somewhere, moderate confidence that Northbridge is exposed, and low confidence about which actor conducted the activity. Do not transfer confidence from one claim to another.

Calibrate judgments over time

Calibration asks whether judgments made with similar likelihood terms prove accurate at roughly the expected rate. If events assessed as “likely” occur only rarely, the team’s language or reasoning needs adjustment.

Maintain a judgment log containing:

  • the exact proposition;
  • assessment date and relevant time horizon;
  • likelihood term and confidence level;
  • supporting evidence, assumptions, and alternatives;
  • expected resolution date or review trigger;
  • later outcome, when observable;
  • reason for any update;
  • lessons for sources, methods, or terminology.

Calibration requires care. Many intelligence questions never resolve cleanly, and absence of a later observation may reflect limited visibility rather than a false judgment. Review only outcomes that can be evaluated responsibly, and distinguish the quality of the original process from luck.

A well-reasoned assessment can lead to an unexpected outcome. A poorly reasoned guess can be correct. Evaluate both outcome accuracy and process quality.

Communicate disagreement honestly

Analytic teams do not always agree. Consensus can improve clarity, but forced consensus hides useful uncertainty.

If disagreement is material to the decision:

  • state the majority or lead judgment;
  • describe the credible alternative view;
  • explain which evidence or assumption drives the difference;
  • identify what new information could resolve it;
  • avoid implying numerical precision based only on a vote.

For example:

The team assesses the activity was likely coordinated, with moderate confidence. One analyst judges an authorized deployment to be roughly as likely because the same administration tool appears in approved maintenance. Verification with the service owner would be decisive.

This communicates structured disagreement rather than presenting indecision. Minor differences that do not affect the consumer’s choice need not dominate the product.

Update judgments without hiding change

An intelligence assessment is time-bounded. New evidence may change likelihood, confidence, scope, or implications. An update should state:

  • what changed;
  • which new evidence caused the change;
  • whether the previous judgment was reasonable given evidence available at the time;
  • what remains unchanged;
  • which decision implications now differ.

Example:

Updated judgment: We now assess the workstation events were very likely coordinated malicious activity, with high confidence, raised from likely with moderate confidence. Memory analysis recovered matching payload behavior on two hosts, and the software owner confirmed no authorized deployment. Credential-access scope remains uncertain because coverage is incomplete.

Changing a judgment is not analytic failure. Failing to change after material evidence arrives is.

Avoid false precision and certainty theater

Do not assign an exact percentage merely to make an assessment look scientific. A judgment of 73% is defensible only if a method supports that precision. Otherwise, a defined verbal range with a clear rationale is more honest.

Also avoid:

  • using confirmed for an inferred conclusion;
  • raising confidence because a senior person agrees;
  • lowering confidence merely because an outcome would be surprising;
  • presenting a single-source claim as high confidence because the source is famous;
  • attaching one confidence label to an entire report containing many propositions;
  • replacing analysis with “more research is needed” without naming the decisive gap;
  • expressing likelihood about an undefined event or time horizon.

Every estimative statement needs a clear proposition. “Ransomware is likely” is incomplete. Likely to occur where, to whom, by when, and under which conditions?

Worked example: assess likelihood and confidence separately

Northbridge is evaluating whether observed activity includes credential access.

Available evidence:

  • two affected hosts launched a tool capable of reading browser-stored credentials;
  • endpoint records show access to browser profile directories;
  • no confirmed credential material was recovered;
  • identity telemetry shows no clearly related sign-in anomaly;
  • visibility is incomplete for one legacy authentication service;
  • benign troubleshooting software can access some of the same directories.

A poor statement is:

We have low confidence that credentials were stolen.

It leaves the probability unclear and overstates what “stolen” means.

A better assessment is:

We assess credential access was roughly as likely as not, with low confidence. Tool execution and browser-directory access are consistent with credential collection, but neither confirms that credential material was obtained. Benign software can produce part of the pattern, and identity visibility is incomplete. Recovery of browser database access, matching memory artifacts, or subsequent account use would increase the likelihood.

The judgment identifies the proposition, likelihood, confidence, evidence, alternative explanation, gap, and change conditions.

Use uncertainty to support decisions

Consumers do not always need the most likely explanation alone. They may need to understand consequences across plausible outcomes.

If an outcome has moderate likelihood but severe impact, a reversible precaution may be justified. If an outcome is very likely but low impact, monitoring may be proportionate. The analyst should communicate:

  • assessed likelihood;
  • confidence and its basis;
  • plausible consequences;
  • warning indicators;
  • time available to decide;
  • reversible and irreversible options;
  • the cost of acting and not acting.

The decision owner combines the intelligence judgment with risk tolerance, operational constraints, legal obligations, and business priorities. Analysts should not manipulate likelihood language to force a preferred action.

Key takeaways
  • Likelihood describes the assessed probability of a proposition; confidence describes the strength of its evidentiary and reasoning foundation.
  • Use a defined estimative vocabulary consistently and explain its meaning to consumers.
  • State confidence for specific judgments and identify the factors that limit it.
  • Do not manufacture numerical precision when the method does not support it.
  • Preserve material analytic disagreement and explain what drives it.
  • Track judgments and outcomes to improve calibration without confusing good process with good luck.
  • Update assessments transparently when evidence changes.

Analyst habit: Write the likelihood judgment first, then add confidence and its rationale. If you cannot state the proposition and time horizon clearly, the assessment is not ready for an estimative term.