What Is Metadata?

Last updated: August 2026

Metadata is structured data about other data: it describes properties of a file or piece of information — such as title, author, creation date, format or keywords — without being the content itself. In organizations, metadata makes digital content findable, governable and automatable; this becomes most tangible with images (EXIF, IPTC, XMP) and product information in DAM and PIM systems.

What is metadata in simple terms?

Put simply, metadata is the profile card of a piece of information. The photo is the content — capture date, camera model, usage rights and keywords are its metadata. The document is the content — author, version and approval status are its metadata. The term derives from the Greek “meta” (“about”): data about data. Whether something is metadata or content depends on the perspective: for a library, the book title describes the book; for a title database, that same title is the content. This descriptive role makes metadata the foundation of every search, sort and automation — no system can find, filter or connect content without it.

What is the difference between data and metadata?

Data is the content itself; metadata describes that content. An email shows the difference cleanly: the message text is the data — sender, recipient, timestamp and subject line are metadata. The difference is functional, not technical: metadata is typically small, highly structured and machine-readable, while the described content can be arbitrarily large and unstructured. That is why systems prefer working with metadata — searching, filtering, sorting and permission checks run vastly faster on structured descriptions than on the content itself.

One adjacent concept deserves a boundary line: meta tags in SEO (title, meta description, robots directives) are metadata of a web page — describing it for search engines and browsers. They are a narrow special case of the concept, not the subject of this page.

What are examples of metadata?

Metadata lives in practically every digital file — examples range from photos to product records:

  • Photo: capture date, camera model, exposure, GPS position (EXIF); description, creator, usage rights (IPTC/XMP); keywords.
  • Document: author, title, creation and modification dates, version, revision history.
  • Email: sender, recipient, timestamp, subject, mail headers.
  • Music/video: title, artist, album, duration, codec, subtitle tracks.
  • Web page: title tag, meta description, canonical and robots directives.
  • Product record: article number, classification, approval status, translation status, channel assignment.

The pattern is always the same: a structured layer describes the actual content — and it is this layer that systems use for search, order and process control.

What types of metadata are there?

Four types have become established in the literature, and they also structure the field setup in DAM and PIM systems:

TypeWhat it describesExamplesTypical benefit
Descriptive metadatacontent and meaningtitle, description, keywords, categoriessearch and findability
Structural metadatastructure and relationshipschapter order, page mapping, variant and file relationsnavigation, versions, context
Administrative metadatamanagement and rightscreator, license, usage rights, expiry date, approval statuslegal certainty, workflows
Technical metadatafile and creation propertiesformat, resolution, color space, codec, camera dataprocessing, output, quality checks

The boundaries are fluid — a keyword can be descriptive in intent and workflow-controlling in use. In practice, the academic classification matters less than the question of which fields your process needs and who maintains them.

What is image metadata (EXIF, IPTC, XMP)?

With images, metadata travels inside the file — via three established standards: EXIF is written by the camera itself (capture time, exposure, lens, GPS where enabled). IPTC originates in the press world and carries descriptive and legal information (caption, creator, copyright notice). XMP is the XML-based framework developed by Adobe that absorbs both worlds and remains extensible — today’s common way to embed metadata in image, PDF and video files. Because the information sits inside the file, it survives transport between systems: a DAM extracts it on import and makes it searchable. One caveat for publishing: many platforms strip embedded metadata on upload — for internal organization, file-embedded maintenance remains the standard nonetheless.

What is document metadata (PDF, Office)?

Office documents and PDFs also carry an invisible profile: author, company, creation and modification dates, editing time, comments and revision history. This cuts both ways: internally, document management and DAM systems use these fields for versioning, search and compliance — externally, a proposal sent to a customer should not expose internal comments or revision histories. Checking and cleaning document metadata before external delivery is therefore basic document hygiene; Office and PDF tools ship inspection and clean-up functions for exactly this.

How do you read, edit or remove metadata?

For a quick look, on-board tools suffice: Windows shows metadata in file properties, macOS in the info dialog and Preview; image programs and PDF readers have their own info panels. For full depth, specialist tools take over — in imaging, ExifTool is the de facto standard for reading every embedded field. And anyone working with files at scale leaves extraction to the system: a DAM reads embedded metadata automatically on import and exposes it as searchable fields. Editing and removal run along the same routes — file properties and image tools rewrite individual fields, clean-up functions strip sensitive entries before files leave the company.

What is metadata management?

Metadata management is the systematic answer to a simple observation: metadata is only as useful as its upkeep. It comprises four building blocks: a schema (which fields exist, which are mandatory, which values are allowed), a taxonomy or controlled vocabularies (consistent terms instead of keyword sprawl), processes (who maintains what, when, with which quality checks) and governance (who may change the schema, how legacy content is migrated). Without these rules, the familiar pattern emerges: everyone tags differently, search finds nothing, and the expensive new system degrades into a file dump. With them, metadata becomes reliable infrastructure for search, automation and reporting.

Why is metadata important?

Because it is the lever that turns stored content into usable content. Metadata does four things: it makes content findable (search and filters operate on metadata, not pixels), governable (rights and expiry dates prevent expired visuals from being reused), connectable (the product image hangs on the right article) and automatable (workflows, syndication and AI processes act on structured fields). The reverse also holds: missing or wrong metadata — not the software — is the most common reason content systems disappoint in daily use.

How does a DAM system use metadata?

In a DAM system, metadata is the operating system of the media library: on import, the DAM extracts embedded fields (EXIF, IPTC, XMP), adds technical properties and creates administrative fields such as rights, license periods and approval status. Descriptive metadata — keywords, categories, descriptions — is added through tagging: the editorial or AI-assisted process that makes content searchable and filterable. How keywords get onto an asset, and what good tagging looks like, is covered on our page What is tagging? — keywords being a subset of metadata, not its synonym. On this foundation, the DAM drives search, collections, rights alerts and channel-appropriate delivery. What to look for when choosing a system: DAM software comparison.

What is metadata in a PIM system (product attributes vs. metadata)?

In the PIM context, one distinction helps: product attributes describe the product itself (material, size, performance, color) — they are the content that later appears on data sheets and product pages. Metadata in the narrower sense describes the data record and its state: approval status, translation status, quality scores, channel assignments, change history. A PIM system needs both layers: attributes make the product describable, metadata makes the maintenance process controllable — for instance when only complete, approved records may be syndicated to a channel. Separating the two cleanly lets you measure data quality and automate workflows instead of tracking states in side spreadsheets.

How does AI-based metadata generation work?

AI generates metadata from the content itself: image recognition suggests keywords, objects and scenes, OCR makes documents searchable, and language models generate descriptions or map content into existing taxonomies. The value lies in scale: tens of thousands of legacy assets can be pre-tagged by machine instead of manually over years. The limits lie in specificity — AI recognizes standard motifs well, while company- and product-specific terms require a maintained vocabulary and editorial control. The proven setup: AI suggests, rules and people decide — within the metadata schema, not around it.

Where does metadata fit in the PIM/PXM/DAM landscape?

Metadata is not a system but the descriptive layer that connects all systems of product and media data management: in the DAM system it makes media findable and legally safe to use, in the PIM system it drives data quality and approvals, and in PXM processes it decides which content plays out in which channel. Media-neutral, well-described data is the precondition for the same content reaching shop, print and marketplace automatically.

Metadata in DAM and PIM with OMN

In OMN, the PIM/DAM platform by apollon, metadata runs end to end: the Digital Asset Management extracts embedded standards such as EXIF, IPTC and XMP on import, manages rights, deadlines and approvals, and makes the library searchable via configurable metadata schemas; AI Tagging supports keywording with AI suggestions within your vocabulary. On the PIM side, status, quality and channel metadata govern product data maintenance through to syndication. The result is one shared, well-described data foundation for all channels — backed by more than 25 years of apollon experience in product data and media processes. Details: OMN Digital Asset Management — or see metadata maintenance and AI support in a free demo.

FAQ — Frequently asked questions about metadata

What does metadata reveal about us (privacy/GDPR)?

More than many expect: GPS data in photos reveals locations, document metadata exposes internal editors and processes — and messengers generate connection metadata (who communicated with whom, and when) that is meaningful even without message content. For companies this means: metadata can constitute personal data under the GDPR — it belongs in privacy concepts and should be reviewed before external release.

Are keywords and metadata the same thing?

No — keywords are a subset: descriptive metadata assigned through tagging. Metadata additionally covers structural, administrative and technical fields that are never “tagged” but captured automatically or set by workflows.

What is a metadata schema?

The binding field structure of a content repository: which metadata fields exist, which are mandatory, which data types and value lists apply. Standards such as Dublin Core provide blueprints; in DAM and PIM projects the schema is usually defined company-specifically.

Does metadata survive export or sending?

Not necessarily. Standards embedded in the file (EXIF, IPTC, XMP) travel with it unless a system strips them — which many social platforms and mail compressors do. System fields of a DAM or PIM (status, rights, relations) live in the database and only leave the system through defined interfaces or export formats.