MISP Data Model: Events, Objects, Attributes and Galaxies
Learn the MISP data model for indicators, events, objects, threat actors, campaigns, sightings, sharing controls, correlation, and lifecycle decisions.
MISP is often described as a threat intelligence platform or an indicator-sharing system. Both descriptions are incomplete. Its real value comes from a data model that can keep a domain, file hash, malware sample, threat actor, campaign, sighting, analyst judgment, and sharing restriction connected without pretending they are the same kind of thing.
That distinction matters because the high-impression search for a “MISP threat intelligence platform database schema” usually reflects a practical question: where do indicators, threat actors, and campaigns belong, and how should they relate? The answer is not a single database table. The public MISP standard is a JSON exchange model centered on an event, with attributes and objects for evidence, taxonomies and galaxies for classification and knowledge, relationships and sightings for context, and distribution controls for safe sharing.
This guide explains that model from the analyst’s perspective. It does not document private database tables or encourage direct database integration. Instead, it shows how to choose the right MISP structure, preserve provenance and uncertainty, avoid common modeling failures, and design exports that downstream detection and response systems can use safely.
The MISP Data Model Is Not the Internal Database Schema
Start with the boundary between an exchange model and an implementation schema.
The official MISP core format defines JSON structures and their semantics so MISP instances and other tools can exchange events and threat information. It explains fields such as UUIDs, timestamps, distribution settings, attributes, objects, galaxies, sightings, organizations, tags, and relationships. This is the contract an integration can reason about.
The application’s relational database is an implementation detail. Tables, indexes, migrations, caches, and internal joins support the software, but they are not the stable intelligence contract. Reading or writing those tables directly can bypass validation, access control, publication state, synchronization behavior, and application logic. It also couples an integration to a particular release.
Use the supported API, PyMISP, or documented import and export formats for automation. Treat UUIDs as the portable identity of exchanged elements; local numeric IDs belong to a specific instance. Preserve timestamps, creator organization, source references, distribution, and object-template versions when moving data. Those fields are not administrative clutter. They are what let another system decide whether two records are the same claim, an update, or unrelated local copies.
A useful mental model is:
- MISP standard: the portable meaning and exchange structure;
- MISP API: the supported application boundary for search and change;
- MISP database: the software’s internal persistence layer;
- your CTI model: the organizational rules that decide what evidence becomes an event, attribute, object, tag, galaxy reference, or relationship.
The last layer is the one teams most often neglect. A technically valid event can still be analytically weak, impossible to maintain, or unsafe to automate.
Events Are the Exchange Envelope, Not Merely Incident Records
In the core format, an event is the main container for a coherent body of threat information. An event may describe an incident, a malware analysis, a campaign update, a threat-actor assessment, a vulnerability exploitation cluster, or a published research report. The event’s meaning comes from the information assembled inside it, not from a rule that every event equals one incident.
An event should have a defensible reason to exist as one package. Ask:
- What claim or investigation binds these elements together?
- Which organization created the event, and which source produced each claim?
- What is the information date and analysis cutoff?
- Is the event complete enough to publish, or still an unpublished working record?
- Who may receive it, and can any child element be distributed more narrowly?
- What would make the event obsolete, superseded, or materially changed?
Avoid two extremes. A giant evergreen event called “all activity from Actor X” becomes hard to version, distribute, and correct. Hundreds of one-attribute events lose the narrative and relationships that make the evidence intelligible. Choose a boundary that matches the intelligence product: one intrusion, one campaign period, one malware-analysis result, or one bounded assessment.
Event Reports can carry the narrative that explains key judgments, sources, gaps, and implications. They complement structured attributes and objects rather than replacing them. This is important for consumers who need to understand why the data matters, not just ingest values.
Use Attributes for Typed Values and Objects for Structured Evidence
An attribute is the smallest commonly exchanged MISP element: a typed value with category, comment, timestamps, distribution, correlation behavior, and other controls. Domains, IP addresses, URLs, hashes, email addresses, vulnerability identifiers, filenames, and free-text values can all be attributes. The type says what the value is; the category helps explain its operational role.
Do not equate “attribute” with “blockable IOC.” Some attributes are observables, contextual values, identities, references, or analyst notes. Whether an attribute should feed a control depends on the evidence and intended use. The IDS flag is one signal, not a substitute for policy. Consumers still need provenance, confidence, first-seen and last-seen context, false-positive risk, and an understanding of what a match would mean.
A MISP object groups attributes that describe one structured entity or observation. The official object-template standard makes those compositions reusable and versioned. A file object can bind a filename, hashes, size, MIME type, signing information, and other properties. A network-connection object can preserve source, destination, ports, protocol, and timing as one observation instead of scattering them across unrelated values.
Use an object when the relationship among fields is part of the evidence. Use separate attributes when each value can stand independently and consumers do not need to know they were observed together. Preserve the object’s template UUID and version so another instance knows which structure produced it.
Object references express typed relationships between objects, including relationships that cross event boundaries where supported. Name the relationship narrowly—such as contains, communicates-with, downloads, or derived-from—and keep the evidence for that edge. A graph with vague related-to links looks connected while saying very little.
Tags, Taxonomies, and Warninglists Add Policy and Classification
Tags attach machine-readable or human-readable labels to events and attributes. A local free-form tag may support a workflow, but uncontrolled tags quickly fragment into spelling variants and private meanings. Taxonomies solve part of that problem by defining governed machine-tag vocabularies.
The official MISP taxonomy format supports classification schemes used for handling, confidence, workflow, threat level, sector, incident classification, and many other purposes. A machine tag follows a structured namespace, predicate, and optional value. The structure lets an automation distinguish a handling restriction from an analytic-confidence statement rather than treating both as arbitrary strings.
Apply tags at the narrowest correct level. Event-level tags describe the whole package. Attribute-level tags express exceptions or claims about one value. If a confidence score applies only to one indicator, placing it on the event can overstate confidence in every other element. If a handling restriction applies to the whole report, repeating it manually on hundreds of attributes creates maintenance risk.
Warninglists address a different problem: values that may be legitimate, common, or too broad for naive correlation, such as public resolver addresses or common infrastructure. A warninglist match is a reason to review context, not automatic proof that the attribute is harmless. Preserve the original evidence and record the decision to suppress, retain, or qualify it.
Design a small required taxonomy profile before importing large feeds. At minimum, decide how the organization expresses provenance, confidence, handling, workflow status, false-positive risk, and retention. If those semantics arrive after ingestion, the team will spend its time repairing ambiguity.
Galaxies and Clusters Represent Threat Actors, Campaigns, Malware, and Other Knowledge
Indicators describe evidence that may be observed. Threat actors and campaigns are higher-level analytic models. MISP galaxies keep those concepts distinct.
The official MISP galaxy format defines a galaxy as a knowledge set containing clusters. A galaxy identifies the kind of knowledge—such as threat actors, malware, tools, techniques, sectors, or another collection—while each cluster represents one concept with a stable UUID, value, description, metadata, and potentially relationships to other clusters.
Attach a galaxy cluster only when the event or attribute supports that association. A report that mentions an actor as an alternative hypothesis does not justify tagging the whole event as activity by that actor. Record whether the source made the claim, whether your team accepts it, and with what confidence. Threat-actor names are not universal identities; different producers can use different names for overlapping activity sets. Our guide to cyber threat attribution explains why those boundaries matter.
Campaign modeling should answer a bounded question. Define the activity set, time range, inclusion criteria, producer, and material relationships. Do not create a campaign cluster merely because several indicators appeared in the same feed. Conversely, do not duplicate a threat-actor cluster every time a new campaign appears. The actor is a model of an operator or activity set; the campaign is a bounded body of activity linked by evidence.
Galaxy clusters add reusable knowledge, but they should never erase the underlying observations. Keep the domains, files, procedures, incidents, and source references that support the association. Otherwise, consumers inherit a confident name without the evidence needed to evaluate it.
Correlation Finds Shared Values; Relationships Explain Why They Matter
MISP can correlate matching attribute values across events. That is useful for discovery: a domain in a new incident may also appear in older malware analysis or partner reporting. Correlation is not attribution. Shared public infrastructure, reused tools, copied reports, scanners, sinkholes, and enrichment artifacts can all create matches that do not prove a common operator.
Use explicit relationships when the connection itself is an analytic claim. Keep the source, time range, and confidence for the edge. If two events are related because one supersedes the other, say so. If a file contacted a domain during sandbox execution, model that observation rather than implying that every file associated with the campaign used the domain.
Analyst data can add notes, opinions, and typed relationships without destructively rewriting another producer’s event. This supports collaboration while preserving who asserted what. Proposals serve a different workflow: they let a recipient suggest additions or changes that the event owner can accept or reject.
This separation mirrors good threat intelligence data-pipeline design: raw observations, source claims, analyst assessments, and downstream decisions should remain traceable even when they refer to the same value.
Sightings and Decaying Models Turn Static Indicators Into Time-Bounded Evidence
An indicator does not remain equally useful forever. Infrastructure is reassigned, phishing domains disappear, malware hashes stop circulating, and benign services are reused by many actors. MISP sightings and decaying models help represent that changing relevance.
A sighting records feedback about an attribute: it was observed, it was assessed as a false positive, or another supported sighting type applies. Sightings add time and operational evidence without requiring every consumer to edit the original indicator. They can also reveal continuing activity after the producing organization expected an indicator to go quiet. Because sightings may expose who observed what, their sharing and anonymization settings require deliberate governance.
The official MISP decaying-model guide describes models that calculate a current relevance score from a base score and time-decay behavior. Models can use taxonomies, sightings, lifetimes, thresholds, and type mappings. The same indicator may decay differently for different uses; a domain may stop being suitable for blocking while remaining relevant to retrospective hunting or investigation.
Do not use decay as automatic truth deletion. A decayed indicator can remain historically important. Separate active-control eligibility from retention of evidence. Export policies should state whether they exclude decayed values, include scores, honor false-positive sightings, or retain historical context for analysts.
Distribution and Sharing Groups Are Part of the Data Model
Threat intelligence is useful only when it reaches an authorized consumer with its restrictions intact. MISP models distribution at the event and, where appropriate, child-element level. The MISP sharing guide documents five distribution choices: your organization only, this community only, connected communities, all communities, and a sharing group.
Publication and distribution are different decisions. An unpublished event remains local to the instance within its applicable distribution rules; publication makes eligible data available for synchronization. Publishing an analytically unfinished event can spread errors. Failing to republish a corrected event can leave partners with stale claims.
Sharing groups are reusable access-control lists spanning selected organizations and instances. Use them when the standard community levels cannot express the intended audience. Keep their purpose and membership governed; a technically valid group can still be wrong for the source agreement or legal basis.
Child elements should normally inherit event distribution unless a specific restriction requires narrower sharing. More restrictive ancestors constrain visibility. That prevents a broadly marked attribute from escaping a narrowly distributed event, but teams should test the actual push, pull, export, and API behavior they rely on.
Handling labels such as TLP communicate the source’s intended sharing boundary, while MISP distribution settings enforce platform visibility. They complement each other and should not be treated as interchangeable. For the wider decision process, see Threat Intelligence Sharing: What to Share, With Whom, and How.
Worked Example: Model a Phishing Campaign Without Flattening the Evidence
Imagine that three organizations report credential-phishing activity targeting the same sector. The reports contain look-alike domains, email subjects, redirect URLs, a reverse-proxy kit, authentication logs, and a vendor claim linking the activity to a named actor.
A defensible MISP model could use:
- One bounded event for the campaign period, with a clear information cutoff, producing organization, source references, distribution, and an Event Report summarizing key judgments and gaps.
- Email objects for messages whose sender, recipient pattern, subject, attachment, and authentication results must remain together.
- Domain, URL, and IP attributes for independently searchable observables, each with first-seen and last-seen context, source comments, IDS eligibility, and false-positive considerations.
- A file or software object for the reverse-proxy kit when several hashes, filenames, or properties describe the same analyzed artifact.
- Object references that state which email contained which URL, which redirect led to which host, and which infrastructure served the kit.
- Taxonomy tags for confidence, handling, workflow, and threat level at the correct scope.
- A campaign galaxy cluster if the bounded activity deserves a reusable campaign identity.
- A threat-actor galaxy cluster only if the attribution claim is represented with its producer, confidence, alternatives, and supporting evidence. A vendor’s actor name should not silently become your organization’s conclusion.
- Sightings when recipients observe the domains or techniques, including false-positive feedback where appropriate.
- Distribution controls that reflect the most restrictive source agreement and allow safe synchronization to the intended community.
The resulting event supports several decisions without collapsing them. A SOC can export eligible current indicators. A hunter can use the behavioral pattern after individual domains decay. An analyst can compare the campaign with other activity. A sharing partner can evaluate the attribution instead of receiving only a name.
Common MISP Modeling Failures and Their Consequences
One attribute equals one conclusion. A matching IP address does not prove compromise, campaign membership, or actor identity. Write the claim a match actually supports.
Everything is marked for IDS export. Context values, shared infrastructure, stale observables, and low-confidence enrichment can flood controls with false positives. Define eligibility and review it over time.
Free-text tags replace governed vocabularies. Inconsistent tags break search, automation, confidence handling, and sharing policy. Use maintained taxonomies for concepts that machines must interpret.
Actor names replace evidence. A galaxy tag is useful for retrieval, but it cannot carry the full attribution argument. Preserve the supporting activity and source claim.
Objects are flattened during integration. Exporting every field as an independent indicator loses which properties were observed together. Map object semantics explicitly and document unavoidable loss.
Correlation is treated as causation. Shared values generate leads, not final judgments. Validate infrastructure ownership, time, source independence, and alternative explanations.
Distribution is applied after enrichment. Imported data can inherit restrictions from its source. Enrichment and fusion must not silently broaden who can receive it.
Corrections do not propagate. Changing an event can unpublish it until it is published again. Test how retractions, deletions, replacements, and false-positive updates reach synchronized instances and downstream tools.
Direct database access becomes an integration. It bypasses the semantic and security boundary provided by the API. The apparent shortcut becomes a long-term migration and integrity risk.
Map MISP to Downstream Systems by Decision, Not by Field Count
A SIEM, EDR, firewall, case platform, graph database, and reporting system do not need identical slices of a MISP event. Start with the consumer decision.
A detection control may need only active, high-confidence, IDS-eligible values with clear expiration, provenance, and response guidance. A case system may need the complete event narrative and analyst notes. A graph workflow may need objects, galaxy clusters, and typed relationships. A data warehouse may retain versioned exports for measurement and audit.
Define transformations explicitly:
- which MISP types map to which downstream fields;
- whether object structure and object references survive;
- how UUIDs, source identity, timestamps, and versions are retained;
- how taxonomies and galaxy clusters map to controlled vocabularies;
- how distribution and handling restrictions are enforced;
- which sightings and decay scores affect eligibility;
- how updates, revocations, and false positives propagate;
- what happens when a receiving system cannot express the original semantics.
Do not call a mapping “lossless” merely because every string arrived. Meaning can be lost when a relationship, scope, time bound, or handling rule disappears. Our guides to STIX and TAXII and CTI integration architecture explain how to design those boundaries.
MISP Data-Model Review Checklist
Before publishing or automating a MISP event, verify:
- Boundary: the event represents one coherent investigation, report, or bounded activity set.
- Identity: portable UUIDs are preserved; local IDs are not treated as global identifiers.
- Evidence: attributes have correct types, categories, source context, and time bounds.
- Structure: fields that must remain together use an appropriate versioned object template.
- Relationships: edges state a meaningful verb and retain their evidential basis.
- Classification: required taxonomies are applied at the correct event or attribute scope.
- Knowledge: galaxies distinguish reusable actors, campaigns, malware, and other concepts from the observations that support them.
- Uncertainty: confidence, alternatives, false-positive risk, and gaps are visible near the claims they qualify.
- Lifecycle: sightings, review dates, and decay behavior support the intended detection, hunting, or historical use.
- Sharing: publication, distribution, sharing groups, and handling labels match source authority and audience.
- Integration: exports preserve provenance and restrictions, and any semantic loss is documented.
- Correction: the team can update, revoke, republish, and propagate a material change.
The objective is not to fill every available field. It is to make each operational claim understandable, reviewable, shareable, and reversible. A smaller event with explicit evidence and governance is more valuable than a densely connected event whose consumers cannot tell what is known, inferred, permitted, or still current.
Frequently asked questions
Is the MISP data model the same as its database schema?
No. The public MISP standard defines an exchange model built around JSON events, attributes, objects, tags, galaxies, relationships, and supporting context. MISP's internal relational database implements application storage and can change between releases. Integrations should normally use the supported API and exchange formats rather than coupling directly to database tables.
What is a MISP event?
A MISP event is the main exchange envelope for a coherent body of threat information. It can represent an incident, investigation, report, campaign update, or threat-actor analysis and can contain attributes, objects, event reports, tags, galaxies, sightings, relationships, provenance, and distribution settings.
What is the difference between a MISP attribute and a MISP object?
An attribute expresses one typed value or indicator, such as a domain, IP address, file hash, vulnerability, or text value. An object groups related attributes according to a versioned template, such as the hashes, filename, size, and signing details that describe one file.
How does MISP represent threat actors and campaigns?
MISP galaxies provide vocabularies of richer entities. A galaxy contains clusters representing concepts such as threat actors, malware, tools, techniques, sectors, or other knowledge sets. Clusters can carry metadata and relationships and can be attached to events or attributes without reducing the underlying evidence to a name alone.
Should every MISP attribute be exported as a detection indicator?
No. Attributes can carry context, references, identities, notes, or observables that are not suitable for blocking or alerting. Type, category, IDS eligibility, confidence, false-positive risk, sightings, time bounds, decay, provenance, and the receiving control's purpose should all influence downstream use.
Is MISP the same as STIX and TAXII?
No. MISP is both a sharing platform and a family of exchange formats centered on events and operational collaboration. STIX models cyber threat intelligence as standardized domain objects and relationships, while TAXII transports CTI collections and objects through an API. Translation is possible, but the models are not lossless equivalents.