Skip to content

Migration Factory on Drupal Migrate

The reusable method for moving content from any source (a legacy Drupal site, another CMS, exports, a live JSON:API) into a structured Drupal content model. It composes Drupal core Migrate and migrate_plus (with migrate_tools for execution) and nothing private. Project-specific mappings — which legacy node becomes which target bundle, field names, term lists — are out of scope here; they live in the owning project's migration map and are cited as evidence, never copied into this playbook.

Ownership rule (specification section 8; module-ownership-matrix.md, migrate_plus): Migration Factory composes Drupal Migrate first and does not create private migration infrastructure where upstream covers the need. Migrations are configuration entities (migrate_plus.migration.*, migrate_plus.migration_group.*) exported with the site and applied through a recipe where reusable; seed scripts and one-off SQL are not migrations.

Where Migrate cannot express a step, prove the gap with the discovery record in how-we-build-in-drupal.md section 2 before writing a process plugin, and write a process plugin — never a bespoke importer.


Steps

Each step produces a named artifact. The artifacts are the migration map (specification section 15), grown one column at a time.

1. SOURCE INVENTORY

  • Enumerate every source item: entity type, bundle, count, published state, URL, last changed, owner. For a Drupal source, drush and the source database; for a live site, JSON:API; for files, a manifest.
  • Record counts as OBSERVED with the command that produced them. Counts drive VALIDATE.
  • Output: SOURCE_INVENTORY — one row per source bundle or file class with counts.

2. CLASSIFICATION

Every source item receives exactly one disposition from the 12-term vocabulary (specification section 15):

MIGRATE_AS_IS                    content is right; move it
MIGRATE_AND_NORMALIZE            move it and conform to the target model (labels, formats, structure)
REWRITE                          the content itself is replaced editorially; migration carries the shell and metadata
EXTRACT_INTO_STRUCTURED_FIELDS   facts buried in prose become fields on the target entity
SPLIT                            one source becomes several target entities
MERGE                            several sources become one target entity
MAP_TO_ENTITY_REFERENCE          a source value becomes a reference to an existing or migrated entity
MAP_TO_TAXONOMY                  a source value becomes a term in a target vocabulary
ARCHIVE                          retained unpublished for record, not migrated to a public target
DELETE                           not migrated; no redirect owed
REDIRECT_ONLY                    no content target; the URL is redirected
HUMAN_REVIEW                     disposition cannot be decided mechanically; a person decides before TRANSFORM

A source item may carry a sequence (for example MIGRATE_AND_NORMALIZE done, REWRITE pending); record the sequence, not the average. HUMAN_REVIEW is a real disposition with an owner and a due decision, not a parking lot.

  • Output: the Classification column of the migration map.

3. TARGET CONTENT MODEL

  • The target model is the structured model the product owns (for bluefly.io, specification sections 4 through 7). This playbook does not define it.
  • Confirm the target bundles, fields, vocabularies, media types and display modes exist as exported configuration before any migration runs. Migration does not create the model; a recipe or the site's configuration does.
  • Apply the structured-content rules in drupal-component-dry-standard.md: fields own facts, entities own business objects, taxonomy classifies, Media owns assets, references build relationships, Views own collections, Canvas owns composition.
  • Output: TARGET_MODEL reference — the configuration set (bundle and field machine names) the migrations will write to.

4. FIELD MAPPING

  • For each source bundle → target bundle pair, map every source field to a target field or to DROP with a reason. Reuse existing field storages before creating a new one.
  • Note format changes (text format, date format, plain vs formatted), length limits, cardinality changes, and required-field gaps that need a default or HUMAN_REVIEW.
  • Express each mapping as a process: pipeline in the migration YAML using core and migrate_plus process plugins (get, default_value, callback, skip_on_empty, sub_process, entity_lookup, entity_generate, migration_lookup, format_date, ...). A custom process plugin needs the proven-gap record.
  • Output: the Target Fields column of the migration map plus the migration YAML process: sections.

5. TAXONOMY MAPPING

  • Map every source classification (categories, tags, free-text labels, boolean flags standing in for categories) to a target vocabulary and term. Prefer shared vocabularies; do not create a vocabulary for one page.
  • Term migrations run first and are referenced with migration_lookup; lookups by name use entity_lookup with the vocabulary bundle constrained. Decide whether unknown values create terms (entity_generate) or fail (skip_on_empty / HUMAN_REVIEW), and record the decision.
  • Output: TAXONOMY_MAP — source value → vocabulary → term, with create-or-fail policy.

6. ENTITY-REFERENCE MAPPING

  • Relationships in the target are entity reference fields (Service ↔ Work, Work ↔ Person, Insight ↔ Topic, ...). Map source relationships (references, embedded lists, repeated prose) to target references.
  • Order migrations by dependency: referenced entities before referencing entities (migration_dependencies: required:). Use migration_lookup to resolve ids; never hardcode target ids.
  • Relationships drive Views automatically (related content, reverse references). Do not migrate "related cards" as content; migrate the reference and let Views render it.
  • Output: RELATIONSHIP_MAP — source relationship → target reference field → dependency order.

7. MEDIA MIGRATION

  • Files migrate to file entities, then to media entities of the target media type (Image, Document, Remote Video, Logo); content references media, not files.
  • Deduplicate by checksum or canonical URI before creating media. Carry alt text, caption, credit, license and focal/crop data where the source has them; where alt text is missing, flag for editorial completion (AI-assisted alt text may propose; a person approves).
  • Remote video becomes a Remote Video media entity by URL; do not download it.
  • Output: MEDIA_MAP — source asset class → media type → dedupe key → required editorial completion.

8. REWRITE RULES

  • For REWRITE, MERGE, SPLIT and EXTRACT_INTO_STRUCTURED_FIELDS, state the rule that produces the target: which prose becomes which field, which sources fold into which target, which structure applies (for example Problem → Consequence → Outcome → Method → Proof → Next Step where the product's copy framework requires it).
  • Rewritten copy is editorial work with an owner; the migration carries metadata, relationships and redirects, and leaves the body in a review state (draft or unpublished moderation state), never publishes rewritten text automatically.
  • Do not add facts, numbers or proof that the source does not carry. A fact without a source is HUMAN_REVIEW.
  • Output: REWRITE_RULES — per classification, the transform and the review owner.

9. REDIRECTS

  • Every source URL that will not resolve identically after migration gets a redirect: MIGRATE_* with a changed alias, MERGE, SPLIT, ARCHIVE, REDIRECT_ONLY. DELETE owes none. Migrate the source's existing redirects too.
  • Owner is the redirect module; redirects are content (a redirect migration) or pathauto alias updates, never Twig or web-server rules for content URLs.
  • Output: REDIRECT_MAP — source path → target path → status code → evidence that it resolves.

10. SEO

  • Owner is metatag (+ token) and pathauto. Migrate meta titles, descriptions, social images and canonical URLs where the source has them; otherwise let target defaults and tokens generate them. AI generation (metatag_ai) is ADOPT_WHEN_NEEDED, after real content exists.
  • Alias patterns are configuration; migrated aliases either match the pattern or get a redirect.
  • Output: metatag and alias configuration referenced from the migration map; no per-page SEO in Canvas.

11. TRANSFORM

  • Write the migrations: one migrate_plus.migration_group per source system with shared source configuration; one migration per source bundle → target bundle, in dependency order; source plugins from core or migrate_plus (url with json / xml data parsers, d7_* / d8_* for Drupal sources, embedded_data for fixtures); destinations entity:node, entity:media, entity:taxonomy_term, entity:redirect, entity:file.
  • Run with migrate_tools: drush migrate:status, drush migrate:import <id> --limit=N for a sample first, then full; drush migrate:rollback must work for every migration (design for idempotence: stable source ids, track_changes or high_water_property for incremental runs).
  • Export the migrations as configuration and commit them with the site (or the recipe that owns them). A migration that exists only on one machine is not done.
  • Output: migration configuration under version control; import and rollback logs as evidence.

12. VALIDATE

  • Counts: target counts per bundle equal the sum of the classification columns that produce a target (source inventory reconciles to migrated + archived + deleted + redirect-only + human-review).
  • Integrity: drush migrate:messages empty or explained; no dangling references (entity_usage report where adopted); required fields populated; media has alt text or an open editorial task; every redirect resolves with the intended status.
  • Model: drush config:status clean after export; drush config:import on a fresh environment succeeds; DRY acceptance in drupal-component-dry-standard.md section 91 holds for pages built on the migrated content.
  • Outcome, not mechanism (drupal-standard.md section 17): a route returning 200 is not proof that the page shows the right content; sample real pages.
  • Output: VALIDATION_RECEIPT — counts, messages, reference integrity, redirect check, config-import proof, each tagged OBSERVED with the command.

13. PUBLISH

  • Publish through the site's moderation workflow, not by migration default: migrated content lands in the state its classification requires (MIGRATE_AS_IS may publish; REWRITE lands in draft; ARCHIVE lands unpublished).
  • Canvas pages consume the published structured content; no canonical copy of migrated content is kept in Markdown or Canvas afterwards (specification section 16).
  • Output: publication state per migration, exported moderation configuration.

14. EVIDENCE

  • The migration map (all columns), migration configuration, import and rollback logs, validation receipt and redirect check are the evidence set. Project-specific evidence stays in the project and its Beads; a dated summary goes to ledger/audits/ or ledger/evidence/ when it proves a reusable lesson (specification section 28).
  • Record the migration as Work where the product treats migrations as proof (bluefly.io does: specification section 15).

Out of scope

  • Which legacy content maps to which target bundle, field and term on a specific site. That is the project's migration map (for bluefly.io: the project's CONTENT_ARCHITECTURE.md section 12 and its Beads), cited from the project, not copied here.
  • The target content model itself. Product knowledge, owned by the product document.
  • Site build and page composition after migration: drupal-site-building-standard.md.

Acceptance

A Migration Factory run is complete when every step above has its artifact, the migrations are configuration under version control, rollback works, VALIDATE is OBSERVED on a fresh environment, and no private migration module or seed script remains in the consumer.