3. Infrastructure, Identity, and Relationship Analysis

Identity, Tooling, and Relationship Hypotheses

Analyze accounts, personas, code, tools, language, infrastructure, and operational patterns without turning association into identity or attribution.

In this lesson, you will learn to:

  • Construct and challenge a relationship hypothesis involving identities, tools, code, or operational patterns while separating observation, association, clustering, coordination, and attribution.

Identity, Tooling, and Relationship Hypotheses

Build relationship hypotheses at appropriate confidence levels, identify shared dependencies and false-link risks, and determine when an identity claim is relevant enough to support an operational decision.

Separate observation, association, clustering, and attribution

Identity analysis asks whether observed accounts, personas, tools, code, infrastructure, and operational patterns are related—and at what level. The analyst must distinguish a documented similarity from a claim that the same person, group, organization, or state controlled the activity.

Use an evidentiary ladder:

Level Meaning Example
Observation A source directly records a feature or event Two scripts contain the same uncommon function and error string.
Association Evidence connects two observations in a bounded way Both scripts were delivered through domains provisioned during the same week.
Technical cluster Several related artifacts or behaviors likely share a production or deployment pattern Code, configuration, and infrastructure form Cluster L-7.
Activity cluster Multiple events or incidents likely reflect related operations Cluster L-7 appears in three finance-focused intrusions.
Common operational control Evidence suggests the same operator or coordinated team controlled the activity Timing, protected account evidence, and operational continuity support shared control.
Attribution Activity is linked to a named person, organization, criminal group, company, or state Multi-source evidence supports attribution to a named entity under appropriate review.

Each level adds inference and consequence. Evidence sufficient to create an internal activity cluster may be inadequate for public attribution.

Define the decision relevance

Before pursuing identity, ask what the answer would change:

  • Would it alter containment or remediation?
  • Would it change which behaviors defenders search for?
  • Would it affect victim, partner, or sector notification?
  • Would it trigger legal, policy, insurance, law-enforcement, or diplomatic action?
  • Would it improve forecasting of likely objectives or next actions?
  • Would it change handling or source-protection requirements?

During Project Lantern, immediate containment depends on behavior and scope, not a named actor. Identity analysis becomes relevant later if it helps connect incidents, anticipate recurring methods, coordinate with partners, or support a consequence-aware attribution decision.

Distinguish entity types

Identity evidence may involve:

Entity type Examples Common ambiguity
Technical account Email, hosting, cloud, registrar, repository, forum, or service account Account compromise, resale, sharing, or automation
Persona Alias, profile, handle, claimed role, or public identity Deliberate deception, imitation, recycled names, or multiple users
Tool or codebase Malware, script, framework, configuration, compiler output Public availability, leakage, purchase, copying, or false flags
Infrastructure cluster Domains, addresses, certificates, tenants, or servers Shared services, compromise, outsourcing, or provider automation
Operational cluster Repeated timing, targeting, procedures, or error patterns Common training, tooling, constraints, or imitation
Organization or group Criminal service, contractor, company, military unit, or state body Fluid membership, affiliates, proxies, overlapping mandates, or uncertain control
Individual Named operator, developer, administrator, or facilitator Shared devices, mistaken identity, coercion, or insufficient legal proof

Do not treat these as interchangeable. Linking a tool to an activity does not identify its operator. Linking an account to infrastructure does not prove who controlled the account during the event.

Record precise relationships

Replace “linked to” with a specific relationship:

  • account-created-resource;
  • persona-posted-content;
  • process-used-tool;
  • sample-derived-from-codebase;
  • infrastructure-served-sample;
  • incident-used-infrastructure;
  • activity-shared-configuration;
  • account-authenticated-to-service;
  • source-reported-common-control;
  • cluster-overlaps-with-cluster;
  • organization-publicly-claimed-activity;
  • provider-confirmed-account-control;
  • possibly-coordinated-with.

For each relationship, record:

  • direction;
  • relevant time period;
  • direct observation or inference;
  • source access and provenance;
  • independence;
  • confidence;
  • alternative explanations;
  • handling and disclosure constraints.
Treat names as source-specific labels

Different vendors, governments, researchers, and communities may use different names for overlapping or distinct activity. A name is a source’s analytical label—not a universal identity.

Maintain a name-mapping record:

Field Question
Source label Which source uses the name?
Definition Which incidents, behaviors, infrastructure, or time period does it include?
Basis What evidence or methodology supports the cluster?
Overlap Which parts overlap with other labels?
Difference Which parts are excluded or unique?
Time When was the mapping valid or last reviewed?
Confidence How certain is the overlap?
Restriction May the mapping be shared or cited?

Use wording such as:

Source A’s Cluster Red overlaps with Northbridge Cluster L-7 in two infrastructure elements and one script family, but available evidence does not establish that the clusters are identical.

Avoid:

Cluster Red is the same as L-7.

unless the evidence truly supports equivalence.

Separate similarity from shared control

Two activities may look similar because of:

  • common public tools;
  • copied code or procedures;
  • shared criminal services;
  • common training or documentation;
  • the same target environment;
  • technical constraints;
  • deliberate imitation or false flags;
  • one operator;
  • coordinated operators;
  • incomplete or biased observation.

A similarity becomes more diagnostic when it is uncommon, independently observed, temporally compatible, difficult to copy, and combined with other evidence types.

Feature Often weak alone Stronger when combined with
Same public tool Many actors can use it Uncommon configuration, timing, infrastructure, and operational sequence
Same malware family Can be sold or shared Protected build identifier, account evidence, development continuity, and deployment pattern
Same language Many speakers and deliberate manipulation Consistent working hours, protected identity evidence, and operational context
Same target sector Many actors share incentives Specific victim selection, timing, access pathway, and objective
Same cloud provider Broad shared service Same tenant, protected administrative identifier, and compatible access logs
Same framework technique Very common abstraction Distinct implementation and relationship sequence
Use cluster identifiers that do not imply identity

Create neutral internal identifiers such as Cluster L-7, Activity Set 2026-04, or Infrastructure Group IG-12. The identifier lets analysts organize evidence while preserving uncertainty.

A cluster record should include:

  • defining evidence;
  • inclusion and exclusion rules;
  • first and last observed;
  • related incidents and entities;
  • confidence that included activity belongs together;
  • leading alternatives;
  • overlap with external labels;
  • changes over time;
  • review and split or merge criteria;
  • intended operational use.

Do not name a cluster after a country, organization, or known actor before attribution is supported. Names influence reasoning and can create confirmation bias.

Evaluate claims and responsibility separately

Several propositions may be confused:

  1. A named persona claimed responsibility.
  2. A technical account was used during the activity.
  3. A person controlled that account.
  4. A group directed the person.
  5. An organization authorized the operation.
  6. A state knew of, supported, tolerated, or controlled the organization.

Evidence for one proposition does not automatically establish the next. A public claim may be truthful, exaggerated, mistaken, opportunistic, or deceptive. Technical involvement does not prove command responsibility.

Preserve temporal identity

Accounts and infrastructure change control. Ask:

  • Who or what controlled the entity during the assessed event?
  • What evidence covers that period?
  • Could credentials or access have been transferred, stolen, rented, or shared?
  • Did provider or legal records refer to account creation, payment, login, or content ownership?
  • Does current control differ from historical control?
  • Are time zones and account logs precise enough to establish overlap?

Do not use current profile ownership to infer historical operation without support.

Protect sensitive and personal information

Identity analysis can affect real people and organizations. Apply necessity, authority, minimization, and consequence-aware review.

Before collecting or sharing identity-related material:

  • define the operational purpose;
  • confirm legal, policy, contractual, and privacy authority;
  • collect the minimum sufficient information;
  • distinguish identifiers from verified identity;
  • assess reidentification and human-impact risk;
  • restrict access and onward sharing;
  • preserve source and method protections;
  • establish retention, correction, and deletion procedures;
  • seek appropriate legal, privacy, policy, or leadership review for consequential claims.

Do not publish personal information, accusations, or public attribution merely because fragments are publicly accessible.

Project Lantern clustering example

Northbridge observes:

  • Script S-1 and an external sample share an uncommon error string and configuration layout.
  • The samples contact different domains but use a similar URL path and timing pattern.
  • Both domains were provisioned through the same reseller within one week.
  • The external sample is reported under Vendor Cluster Ember.
  • The tool family is privately sold to several customers.
  • Available evidence does not reveal account ownership or operator identity.

A defensible conclusion is:

Project Lantern activity and part of Vendor Cluster Ember likely share a tooling and provisioning pattern, with moderate confidence. The overlap supports an internal activity-cluster hypothesis and behavior-focused hunting. Because the tool is sold to multiple customers and no protected identity evidence is available, the finding does not establish a common operator or named actor.

Relationship quality checklist

Before raising an identity claim, confirm:

  • The decision relevance and consequence are explicit.
  • Observation, association, technical cluster, activity cluster, common control, and attribution are distinguished.
  • Entity types and relationship types are precise.
  • Names are treated as source-specific labels with documented definitions and overlap.
  • Similarity is evaluated against common use, copying, service sharing, deception, and observation bias.
  • Every relationship is temporally compatible and provenance-rich.
  • Claims about accounts, people, groups, organizations, and states remain separate.
  • Neutral cluster identifiers preserve uncertainty.
  • Personal, source, legal, policy, and disclosure risks are controlled.
  • The evidence threshold matches the intended operational or public use.
Key takeaways
  • Identity analysis is an evidentiary ladder; each step from observation to attribution adds inference and consequence.
  • Define what an identity answer would change before investing in it.
  • Distinguish accounts, personas, tools, infrastructure, activity clusters, organizations, and individuals.
  • Replace vague association language with precise, time-bounded relationships.
  • Similar behavior may result from shared tools, services, constraints, copying, or deception—not only common control.
  • Treat external actor names as source-specific analytical labels.
  • Use neutral internal clusters until stronger identity evidence is available.
  • Apply heightened handling and review to claims that could harm people, organizations, partners, or policy interests.

Analyst habit: When writing an identity claim, identify the exact rung of the ladder it occupies and remove every word that implies a higher rung than the evidence supports.

Test identity and tooling hypotheses without overclaiming

A defensible identity or tooling hypothesis states exactly which relationship is being assessed, during which period, and for which decision. It then compares that explanation with plausible alternatives rather than treating similarity as proof.

For Project Lantern, analysts might test:

Project Lantern and the activity described as Vendor Cluster Ember were likely produced through a shared tooling and provisioning process, but available evidence is insufficient to establish the same operator.

This proposition separates a supportable technical relationship from the stronger claim of common operational control.

Decompose compound identity claims

A statement such as “Cluster Ember conducted Project Lantern” contains several propositions:

  1. The samples belong to the same tool or code lineage.
  2. The infrastructure reflects a related provisioning process.
  3. The incidents form one activity cluster.
  4. The same operator controlled both activities.
  5. Vendor Cluster Ember accurately defines that operator.
  6. A named person, organization, or state is responsible.

Evaluate each proposition separately. Confidence may be high for code lineage, moderate for activity clustering, and low for operator identity.

Proposition Evidence needed Common alternative
Shared code lineage Uncommon implementation overlap, development artifacts, version relationships Public or commercially shared code
Shared provisioning Account, configuration, timing, or protected administrative overlap Common reseller or automation
Related activity Compatible behavior, targets, infrastructure, timing, and operational sequence Copying, coincidence, or shared service
Common operator Continuity across protected accounts, decisions, access, or operational control Affiliates, customers, contractors, or tool sharing
Named attribution Multi-source identity and responsibility evidence with appropriate review Misidentification, deception, or label mismatch
Create competing hypotheses

For the Project Lantern overlap:

  • H1 — Same operator: One operator or tightly coordinated team conducted both activities.
  • H2 — Shared tooling customer: Different operators used the same privately sold tool and provisioning service.
  • H3 — Copy or derivation: One party copied code, configuration, or public reporting from another.
  • H4 — Common upstream provider: A developer, contractor, affiliate, or infrastructure service supplied both operations.
  • H5 — Analytical or dataset error: Parsing, sample contamination, label mismatch, or dependent reporting created the apparent overlap.

A useful hypothesis set includes explanations at comparable levels. “Same nation-state” should not compete with the vague alternative “coincidence.”

Inventory independent evidence types

Identity judgments strengthen when several substantially independent evidence types converge:

Evidence family Examples Limitation
Code and build Function lineage, compiler artifacts, configuration format, protected build identifiers Tools can be sold, leaked, copied, or deliberately altered
Infrastructure Tenant identifiers, account continuity, provisioning pattern, service configuration Shared providers and compromised accounts create false links
Behavior Action sequence, operational security, error handling, persistence, targeting Common tradecraft and imitation reduce uniqueness
Temporal pattern Working periods, deployment cadence, resource lifetime, response to disruption Automation, global teams, and incomplete observation distort timing
Victimology Sector, geography, organization type, access pathway, objective Popular targets attract unrelated operators
Human or account evidence Provider records, authenticated access, payment, communication, repository activity Sensitive, incomplete, transferable, and consequence-heavy
Partner reporting Direct incident access, protected investigation, provider confirmation May share underlying sources or definitions
Claims and public behavior Statements, aliases, recruitment, sales, or leak-site activity Deception, opportunism, and impersonation are common

Count evidence families only after tracing dependencies. A report, blog, and vendor feed may all repeat one protected observation.

Evaluate code and tooling carefully

Tool similarity can be described at several levels:

  • same exact file;
  • modified build of the same codebase;
  • shared library or component;
  • similar configuration schema;
  • compatible protocol or command set;
  • functionally similar implementation;
  • same public framework;
  • common compiler or development environment.

The more abstract the similarity, the less identity value it usually provides.

Ask:

  • Is the feature uncommon among comparable tools?
  • Could it come from a public dependency or default template?
  • Is the sample complete and provenance-rich?
  • Could one sample have been contaminated or mislabeled?
  • Does version order support derivation?
  • Do protected build, account, or deployment features accompany the code overlap?
  • Does behavior in the incident match what the tool could do, or merely what it contains?

A capability present in code is not proof it was used during the operation.

Test temporal and operational continuity

Common control becomes more plausible when evidence shows continuity across decisions that are difficult to explain through shared tools alone:

  • coordinated infrastructure replacement after disruption;
  • consistent protected account access;
  • recurring victim selection tied to one objective;
  • stable operational procedures and error recovery;
  • development changes followed by compatible deployment;
  • continuity across private communications or provider records;
  • distinctive sequencing across several incidents.

Yet continuity can also come from a shared service or playbook. Compare the timing with provider automation, customer onboarding, affiliate activity, and public disclosure.

Use disconfirming tests

For every favored identity hypothesis, state evidence that would weaken it.

Hypothesis Would strengthen Would weaken
Same operator Protected account continuity, uncommon shared implementation, coordinated deployment and response Evidence of separate customers, overlapping but incompatible operations, or independently transferred tooling
Shared tooling customer Provider or developer evidence of multiple customers, divergent infrastructure and victim choices Protected operational account shared across incidents
Copied implementation Public release predates overlap; version history shows derivation Private feature appears independently before any plausible access
Common upstream provider Shared administrative service, provisioning account, or contractor evidence Direct evidence that each operator built and managed resources independently
Dataset error Broken lineage, mislabeled sample, parser defect, or circular reporting Independent preserved observations reproduce the relationship

Do not phrase change conditions so vaguely that no evidence could overturn the judgment.

Avoid behavioral stereotypes

Working hours, language settings, keyboard layouts, cultural references, and holidays may add context, but they rarely establish identity alone. They can reflect:

  • victim or system configuration;
  • remote infrastructure location;
  • globally distributed teams;
  • automation schedules;
  • deliberate deception;
  • copied code;
  • analyst selection bias;
  • incomplete observation.

Use these features only with provenance, appropriate comparison data, and stronger evidence. Avoid inferring nationality, ethnicity, or state control from cultural or linguistic fragments.

Assess confidence at the relationship level

A relationship record should explain:

  • the exact proposition;
  • likelihood and confidence;
  • strongest supporting evidence;
  • strongest alternative;
  • source independence;
  • temporal compatibility;
  • evidence gaps;
  • consequence and intended use;
  • change conditions.

Example:

We assess Project Lantern and part of Vendor Cluster Ember likely share a tooling and provisioning process, with moderate confidence. The activities use a related codebase, an uncommon response path, and temporally compatible provisioning patterns observed through two independent sources. Confidence is limited because the tool is sold to multiple customers and no protected account evidence establishes common control. Use the relationship for behavior-centered hunting and partner coordination, not named attribution.

Match the threshold to the use
Intended use Suitable evidentiary level
Enrichment A documented association with visible uncertainty
Internal clustering Multiple compatible observations with inclusion and exclusion rules
Hunting and detection Behavior or technical relationships that can be tested independently
Partner coordination Sanitized evidence, confidence, definitions, handling, and feedback request
Strategic forecasting Stable activity patterns and explicit assumptions about continuity
Legal, personnel, policy, or public action Higher-confidence, multi-source evidence plus appropriate authority and review

An internal cluster can be useful without a public name. Attribution should not be pursued merely to make the product feel complete.

Manage cluster evolution

Clusters should change as evidence changes. Define rules to:

  • add an incident or entity;
  • exclude a weak or contradicted relationship;
  • split one cluster into several;
  • merge clusters when stronger evidence supports equivalence;
  • supersede an external-name mapping;
  • expire stale relationships;
  • preserve previous versions and reasons for change.

Example:

Cluster L-7 initially included Hosts A, B, and C because all showed a related administration tool. Package validation later demonstrated that Host C’s activity was authorized and technically distinct. Cluster version 1.2 excludes Host C and records the change rather than rewriting the original record.

Conduct consequence-aware review

Before a consequential identity claim is released, review:

Analytic integrity
  • Are individual propositions and confidence levels distinct?
  • Were alternatives and disconfirming evidence considered?
  • Is source independence understood?
  • Does the wording exceed the evidence rung?
Human and organizational impact
  • Could the claim unfairly implicate a person, company, community, or state?
  • Is personal information necessary and lawfully handled?
  • Could errors cause retaliation, litigation, employment harm, or diplomatic consequence?
Source and operational protection
  • Could disclosure expose access, partners, methods, investigations, or collection gaps?
  • Could the subject adapt or destroy evidence?
Authority and purpose
  • Who is authorized to make and release the claim?
  • Which decision requires it?
  • Is a less identifying cluster statement sufficient?
Worked Project Lantern assessment

The evidence shows:

  • Script S-1 and an external sample share an uncommon configuration layout and error-handling routine.
  • The external sample predates Project Lantern and belongs to a tool reportedly sold to several customers.
  • Incident domains use different providers but share a rare URL path and compatible service behavior.
  • Provisioning occurred during the same week through one reseller.
  • No protected provider, payment, repository, or account evidence connects the operators.
  • Targeting overlaps in finance-related organizations but is not unique.

Assessment:

Project Lantern likely overlaps technically with the activity labeled Vendor Cluster Ember, with moderate confidence. Code lineage and uncommon service behavior support a shared tooling ecosystem. Common operator control is possible but not assessed as likely because the tool has multiple customers and identity-bearing evidence is absent. Northbridge will maintain neutral Cluster L-7 for internal hunting and partner comparison. Any claim of named responsibility requires additional independent evidence and consequence-aware review.

Identity-hypothesis workflow
  1. Define the decision and exact relationship proposition.
  2. Break compound claims into separate rungs.
  3. Generate same-operator, shared-service, copied-tooling, upstream-provider, and error hypotheses.
  4. Inventory evidence by family and trace dependencies.
  5. Test uniqueness, temporal compatibility, and transferability.
  6. Seek disconfirming and identity-bearing evidence proportionate to the decision.
  7. Assign likelihood and confidence to each relationship—not the entire narrative.
  8. Select an operational use that matches the evidence.
  9. Record change conditions and cluster-version rules.
  10. Apply handling and consequence-aware review before disclosure.
Common failure modes
  • Tool-owner equivalence: Whoever used a tool is assumed to have developed or exclusively controlled it.
  • Name inheritance: A vendor label is copied without its definition or uncertainty.
  • Country-by-clock: Working hours or language become nationality claims.
  • Infrastructure ownership leap: Rental or use is treated as legal or organizational ownership.
  • Evidence-family counting: Several dependent reports appear to be multi-source corroboration.
  • Cluster permanence: Inclusion decisions never change after contrary evidence.
  • Attribution as completion: Analysts pursue identity even when it cannot change the decision.
  • Confidence transfer: High confidence in code similarity inflates confidence in common operator or state responsibility.
  • Disclosure mismatch: Internal hunt-level evidence is presented publicly as attribution.
Key takeaways
  • Break identity claims into code lineage, provisioning, activity clustering, common control, and named attribution.
  • Test same-operator explanations against tool sharing, copying, upstream services, and data error.
  • Use multiple independent evidence families and evaluate uniqueness, timing, and transferability.
  • Treat behavioral and cultural features cautiously and avoid identity stereotypes.
  • Assign likelihood and confidence to precise relationships.
  • Match the evidentiary threshold to enrichment, clustering, hunting, sharing, forecasting, or consequential attribution.
  • Version clusters and preserve why entities were added, excluded, split, or merged.
  • Prefer a useful neutral cluster over an unsupported name.

Analyst habit: Ask, If the shared tool were available to several operators, which independent evidence would still connect these activities? That evidence deserves the most weight.