Skip to content

Latest commit

 

History

History

README.md

DPLA Ingest Pipeline — Runbook

All scripts communicate with the ingest EC2 instance via AWS SSM — you never need to SSH in directly.


How to Run an Ingest

Three paths are available depending on how much control you need:

Method When to use
1. Manual via SSM Debugging, one-off runs, or when you need to run a command directly on the box
2. Manual via Python scripts Normal monthly operations — run from your Mac, full visibility, interactive prompts
3. GitHub Actions Hands-off launches — trigger from the GitHub UI; EC2 does the work

1. Manual via SSM

Send commands directly to the EC2 instance:

# Ad-hoc ingest (replace <hub> with the hub short name)
aws ssm send-command \
  --instance-ids i-XXXXX \
  --document-name AWS-RunShellScript \
  --parameters 'commands=["sudo -u ec2-user bash -l -c \"nohup bash /home/ec2-user/ingestion3/scripts/ingest.sh <hub> > /home/ec2-user/data/<hub>-ingest.log 2>&1 </dev/null &\""]' \
  --region us-east-1

Use SSM for one-off commands, quick checks, or when the Python scripts aren't available.

2. Manual via Python scripts

Run from your Mac — see the sections below for each step of the monthly workflow.

python3 hub_preflight.py      # pre-flight
python3 launch_ingest.py bpl  # launch a hub
python3 check_ingest.py bpl   # monitor it

3. GitHub Actions

Two workflows are available under Actions → Launch Hub Ingest / Monthly Hub Batch:

Workflow Purpose File
Launch Hub Ingest Manual single-hub trigger .github/workflows/ingest-hub.yml
Monthly Hub Batch Reads i3.conf, fires all scheduled standard hubs sequentially; special-case hubs (NARA, Smithsonian, Community Webs) are excluded and must be run separately .github/workflows/ingest-monthly.yml

Launch Hub Ingest inputs:

  • hub (required) — short name, e.g. bpl, ohio, nara, smithsonian
  • resume_from (optional) — resume standard hubs from harvest|mapping|enrichment|jsonl|delivery
  • override (optional) — NARA: YYYYMM; Smithsonian: YYYY-MM-DD; Community Webs: YYYYMMDD_HHMMSS

Monthly Hub Batch inputs:

  • month (optional, 1-12) — defaults to current month
  • dry_run — print hub list without launching

Both workflows require the INGEST_INSTANCE_ID secret and OIDC auth scoped to main. The batch fires all hubs as a sequential background job on EC2 — the GHA step completes immediately; Slack notifies as each hub finishes.


Prerequisites

  • Python 3.9+
  • AWS CLI installed and authenticated (~/.aws/credentials)
  • [nara] profile in ~/.aws/credentials (obtain credentials from Dominic)
  • sbt installed locally for batch JAR builds (brew install sbt)
  • ~/.dpla-secrets.env locally with DPLA_API_KEY
  • ~/.dpla-secrets.env on the ingest EC2 with SLACK_BOT_TOKEN and SLACK_CHANNEL

First-Time Setup

python3 onboarding.py

Prompts for your local repo paths (ingestion3, ingestion3-conf, batch-process-dpla-index) and saves them to .env in the ingestion3 root (gitignored). Also verifies your AWS credentials, NARA profile, i3.conf, and DPLA API key.

Re-run at any time. Use --update to change saved paths.


Monthly Workflow

Each month runs in roughly this order. Special-case hubs (Smithsonian, NARA, Community Webs) run their own scripts but fit into the same overall sequence.

0. Monthly hub list     →  pre_ingest_check.py   (list all hubs for the month)
1. Pre-flight checks    →  hub_preflight.py       (per hub, before each ingest)
2. Provider ingests     →  launch_ingest.py       (all standard hubs)
3. Special Cases        →  <hub_name>/launch_hub.py
6. Index rebuild        →  launch_indexer.py
7. Post-index batch     →  post_indexer.py
8. Verify               →  postchecks.py

Standard Hub Ingest

Step 0 — Monthly hub list

python3 pre_ingest_check.py              # current month
python3 pre_ingest_check.py --month 5    # specific month

Lists all hubs scheduled for the month from i3.conf, with harvest types, special-case notes, and on-hold status. Run this at the start of each ingest month to plan the run order.


Step 1 — Pre-flight checks

python3 hub_preflight.py

Prompts the user for the hub endpoint to check. Checks the EC2 instance state, repo freshness on EC2 (ingestion3 + ingestion3-conf), JAR freshness, disk space, and reachability of the hub's harvest endpoint. Starts the instance automatically if it's stopped.

Run for every hub before launching.

Flag options:

python3 hub_preflight.py --endpoint-only   # just check the endpoint
python3 hub_preflight.py --skip-endpoint   # skip the endpoint check
python3 hub_preflight.py --no-start        # don't auto-start EC2 if stopped

Step 2 — Launch ingest

python3 launch_ingest.py

Prompts the user for the hub to run. Runs ingest.sh <hub> on EC2 in the background via SSM. For file-type hubs (Ohio, Georgia, Florida, etc.) it lists available S3 deliveries and asks you to confirm the endpoint before launching.

To resume a failed run partway through:

python3 launch_ingest.py <hub> --resume-from mapping     # or enrichment / jsonl

Step 3 — Monitor

python3 check_ingest.py 
python3 check_ingest.py --watch    # auto-refresh every 30s
python3 check_ingest.py --watch 60 # custom interval

Shows process status, completed stages with record counts, current stage, disk usage, and recent log lines.


Step 4 — Launch indexer

Once all hubs for the month are ingested:

python3 launch_indexer.py

Runs pre-flight checks (no existing cluster, snapshot ages, batch output path, JAR freshness), launches the sparkindexer EMR cluster, monitors it with an hourly Slack heartbeat, then walks you through the Elasticsearch alias swap. Cluster takes 6–9 hours.

python3 launch_indexer.py --cluster-id j-XXXXX   # resume monitoring an existing cluster
python3 launch_indexer.py --alias-swap-only       # skip straight to alias swap
python3 launch_indexer.py --verify-only           # just verify the API count
python3 launch_indexer.py --skip-preflight        # skip pre-flight checks

Step 5 — Post-indexer

Immediately after the alias swap completes:

python3 post_indexer.py

Checks/rebuilds the batch JAR, launches the monthlybatch EMR cluster, and runs four Spark steps in sequence: ParquetDump → JsonlDump → MqReports → Sitemap. On completion it runs hub stats on EC2 and triggers the hub sitemaps GitHub Actions workflow. Cluster takes 3–5 hours.

python3 post_indexer.py --cluster-id j-XXXXX   # resume monitoring
python3 post_indexer.py --skip-preflight        # skip JAR check

Step 6 — Verify

python3 postchecks.py                    # current month
python3 postchecks.py --month 202604     # specific month (YYYYMM)

Reads i3.conf to get the hub list scheduled for that month, checks snapshot dates and record counts in S3, verifies the provider export, hub stats, and sitemaps, and hits the live API for a total record count.


Special Cases

NARA

NARA delivers delta files bimonthly (Feb, Apr, Jun, Aug, Oct, Dec) via their own S3 bucket (s3://ngc-storage01), which requires separate NARA-issued credentials. The process is three steps.

Step 1 — Stage files

python3 nara/copy_nara.py --month 202604

Runs on EC2 via SSM. Downloads ZIPs from ngc-storage01 using the [nara] AWS profile, uploads them to s3://dpla-hub-nara/raw_ingest_files/<YYYYMM>/, then moves them to the EC2 ingest directory so nara-ingest.sh doesn't need to re-download them.

NARA credentials live in ~/.aws/credentials on EC2 under [nara]. If they're expired, email tech@dp.la.

Step 2 — Full Pipeline

python3 nara/launch_nara.py --month 202604

Runs a preflight check that files are staged on EC2, then launches nara-ingest.sh in the background. NARA is a large delta ingest — expect several hours.

Step 3 — Monitor

python3 nara/check_nara.py
python3 nara/check_nara.py --watch

Shows process, completed stages, latest merged harvest record count, disk usage, and recent log lines.


Smithsonian

Smithsonian delivers files to s3://dpla-hub-si bimonthly (Feb, Apr, Jun, Aug, Oct, Dec). It requires a preprocessing step (fix-si.sh) before the standard pipeline and has optional human review checkpoints at each stage.

python3 smithsonian/launch_smithsonian.py

Stages run in order with a confirmation prompt between each:

  1. Download — syncs from s3://dpla-hub-si to EC2
  2. Preprocess — runs fix-si.sh to normalize raw files (checkpoint)
  3. Harvest — runs harvest.sh smithsonian (checkpoint)
  4. Mapping — runs ingest.sh smithsonian --mapping-only (checkpoint)
  5. Pipeline — runs ingest.sh smithsonian --resume-from enrichment
python3 smithsonian/launch_smithsonian.py --auto              # skip checkpoints
python3 smithsonian/launch_smithsonian.py --date 2026-04-03   # specific delivery date
python3 smithsonian/launch_smithsonian.py --start-at mapping  # resume from a stage

Monitor with:

python3 smithsonian/check_status_smithsonian.py
python3 smithsonian/check_status_smithsonian.py --watch

Community Webs

Internet Archive delivers a SQLite .db file directly (not via S3). File recieved to tech@dp.la and downloaded to local machine.

python3 community-webs/launch_cw.py --db ~/Downloads/community-webs.db --full

This:

  1. Uploads the .db to s3://dpla-scratch/community-webs/
  2. Downloads it to EC2 at a temp path
  3. Runs community-webs-ingest.sh --db=<path> --full on EC2: export → harvest → mapping → enrichment → jsonl → S3

Without --full, only export and harvest run. Re-run with --full and --skip-export to continue:

python3 community-webs/launch_cw.py --full --skip-export

Skip flags for resuming after a failure:

--skip-upload    # .db already in s3://dpla-scratch/community-webs/
--skip-export    # ZIP already on EC2, skip straight to harvest

Monitor with:

python3 community-webs/check_cw.py
python3 community-webs/check_cw.py --watch

Slack Notifications

Scripts send notifications to #tech-alerts via the ingest EC2. Configure webhooks in .env:

SLACK_WEBHOOK=           # #tech-alerts
SLACK_TECH_WEBHOOK=      # #tech (hub-complete with @here)
SLACK_ALERT_USER_ID=     # your Slack member ID for @mention on failures

Troubleshooting

Clock skew / InvalidSignatureException from AWS CLI AWS rejects requests when your system clock is off. Fix with:

sudo sntp -sS time.apple.com

Then re-run or use --resume / --cluster-id to pick up where you left off.

Ingest failed mid-run Use --resume-from <stage> in launch_ingest.py to restart from mapping, enrichment, or jsonl without re-harvesting.

Cluster monitoring crashes Your IAM user needs EMR read permissions. Resume with --cluster-id j-XXXXX once permissions are sorted.

Slack notifications not sending Verify ~/.dpla-secrets.env exists on the ingest EC2 with SLACK_BOT_TOKEN and SLACK_CHANNEL set.

sbt assembly fails Make sure sbt is installed (brew install sbt on Mac) and batch-process-dpla-index is cloned at the path saved in your .env.

NARA credentials expired The [nara] profile in ~/.aws/credentials on EC2 uses time-limited keys. Request new ones from NARA and update the profile.