All scripts communicate with the ingest EC2 instance via AWS SSM — you never need to SSH in directly.
Three paths are available depending on how much control you need:
| Method | When to use |
|---|---|
| 1. Manual via SSM | Debugging, one-off runs, or when you need to run a command directly on the box |
| 2. Manual via Python scripts | Normal monthly operations — run from your Mac, full visibility, interactive prompts |
| 3. GitHub Actions | Hands-off launches — trigger from the GitHub UI; EC2 does the work |
Send commands directly to the EC2 instance:
# Ad-hoc ingest (replace <hub> with the hub short name)
aws ssm send-command \
--instance-ids i-XXXXX \
--document-name AWS-RunShellScript \
--parameters 'commands=["sudo -u ec2-user bash -l -c \"nohup bash /home/ec2-user/ingestion3/scripts/ingest.sh <hub> > /home/ec2-user/data/<hub>-ingest.log 2>&1 </dev/null &\""]' \
--region us-east-1Use SSM for one-off commands, quick checks, or when the Python scripts aren't available.
Run from your Mac — see the sections below for each step of the monthly workflow.
python3 hub_preflight.py # pre-flight
python3 launch_ingest.py bpl # launch a hub
python3 check_ingest.py bpl # monitor itTwo workflows are available under Actions → Launch Hub Ingest / Monthly Hub Batch:
| Workflow | Purpose | File |
|---|---|---|
| Launch Hub Ingest | Manual single-hub trigger | .github/workflows/ingest-hub.yml |
| Monthly Hub Batch | Reads i3.conf, fires all scheduled standard hubs sequentially; special-case hubs (NARA, Smithsonian, Community Webs) are excluded and must be run separately | .github/workflows/ingest-monthly.yml |
Launch Hub Ingest inputs:
hub(required) — short name, e.g.bpl,ohio,nara,smithsonianresume_from(optional) — resume standard hubs fromharvest|mapping|enrichment|jsonl|deliveryoverride(optional) — NARA:YYYYMM; Smithsonian:YYYY-MM-DD; Community Webs:YYYYMMDD_HHMMSS
Monthly Hub Batch inputs:
month(optional, 1-12) — defaults to current monthdry_run— print hub list without launching
Both workflows require the INGEST_INSTANCE_ID secret and OIDC auth scoped to main.
The batch fires all hubs as a sequential background job on EC2 — the GHA step completes
immediately; Slack notifies as each hub finishes.
- Python 3.9+
- AWS CLI installed and authenticated (
~/.aws/credentials) [nara]profile in~/.aws/credentials(obtain credentials from Dominic)sbtinstalled locally for batch JAR builds (brew install sbt)~/.dpla-secrets.envlocally withDPLA_API_KEY~/.dpla-secrets.envon the ingest EC2 withSLACK_BOT_TOKENandSLACK_CHANNEL
python3 onboarding.pyPrompts for your local repo paths (ingestion3, ingestion3-conf, batch-process-dpla-index) and saves them to .env in the ingestion3 root (gitignored). Also verifies your AWS credentials, NARA profile, i3.conf, and DPLA API key.
Re-run at any time. Use --update to change saved paths.
Each month runs in roughly this order. Special-case hubs (Smithsonian, NARA, Community Webs) run their own scripts but fit into the same overall sequence.
0. Monthly hub list → pre_ingest_check.py (list all hubs for the month)
1. Pre-flight checks → hub_preflight.py (per hub, before each ingest)
2. Provider ingests → launch_ingest.py (all standard hubs)
3. Special Cases → <hub_name>/launch_hub.py
6. Index rebuild → launch_indexer.py
7. Post-index batch → post_indexer.py
8. Verify → postchecks.py
python3 pre_ingest_check.py # current month
python3 pre_ingest_check.py --month 5 # specific monthLists all hubs scheduled for the month from i3.conf, with harvest types, special-case notes, and on-hold status. Run this at the start of each ingest month to plan the run order.
python3 hub_preflight.pyPrompts the user for the hub endpoint to check. Checks the EC2 instance state, repo freshness on EC2 (ingestion3 + ingestion3-conf), JAR freshness, disk space, and reachability of the hub's harvest endpoint. Starts the instance automatically if it's stopped.
Run for every hub before launching.
Flag options:
python3 hub_preflight.py --endpoint-only # just check the endpoint
python3 hub_preflight.py --skip-endpoint # skip the endpoint check
python3 hub_preflight.py --no-start # don't auto-start EC2 if stoppedpython3 launch_ingest.pyPrompts the user for the hub to run. Runs ingest.sh <hub> on EC2 in the background via SSM. For file-type hubs (Ohio, Georgia, Florida, etc.) it lists available S3 deliveries and asks you to confirm the endpoint before launching.
To resume a failed run partway through:
python3 launch_ingest.py <hub> --resume-from mapping # or enrichment / jsonlpython3 check_ingest.py
python3 check_ingest.py --watch # auto-refresh every 30s
python3 check_ingest.py --watch 60 # custom intervalShows process status, completed stages with record counts, current stage, disk usage, and recent log lines.
Once all hubs for the month are ingested:
python3 launch_indexer.pyRuns pre-flight checks (no existing cluster, snapshot ages, batch output path, JAR freshness), launches the sparkindexer EMR cluster, monitors it with an hourly Slack heartbeat, then walks you through the Elasticsearch alias swap. Cluster takes 6–9 hours.
python3 launch_indexer.py --cluster-id j-XXXXX # resume monitoring an existing cluster
python3 launch_indexer.py --alias-swap-only # skip straight to alias swap
python3 launch_indexer.py --verify-only # just verify the API count
python3 launch_indexer.py --skip-preflight # skip pre-flight checksImmediately after the alias swap completes:
python3 post_indexer.pyChecks/rebuilds the batch JAR, launches the monthlybatch EMR cluster, and runs four Spark steps in sequence: ParquetDump → JsonlDump → MqReports → Sitemap. On completion it runs hub stats on EC2 and triggers the hub sitemaps GitHub Actions workflow. Cluster takes 3–5 hours.
python3 post_indexer.py --cluster-id j-XXXXX # resume monitoring
python3 post_indexer.py --skip-preflight # skip JAR checkpython3 postchecks.py # current month
python3 postchecks.py --month 202604 # specific month (YYYYMM)Reads i3.conf to get the hub list scheduled for that month, checks snapshot dates and record counts in S3, verifies the provider export, hub stats, and sitemaps, and hits the live API for a total record count.
NARA delivers delta files bimonthly (Feb, Apr, Jun, Aug, Oct, Dec) via their own S3 bucket (s3://ngc-storage01), which requires separate NARA-issued credentials. The process is three steps.
Step 1 — Stage files
python3 nara/copy_nara.py --month 202604Runs on EC2 via SSM. Downloads ZIPs from ngc-storage01 using the [nara] AWS profile, uploads them to s3://dpla-hub-nara/raw_ingest_files/<YYYYMM>/, then moves them to the EC2 ingest directory so nara-ingest.sh doesn't need to re-download them.
NARA credentials live in
~/.aws/credentialson EC2 under[nara]. If they're expired, email tech@dp.la.
Step 2 — Full Pipeline
python3 nara/launch_nara.py --month 202604Runs a preflight check that files are staged on EC2, then launches nara-ingest.sh in the background. NARA is a large delta ingest — expect several hours.
Step 3 — Monitor
python3 nara/check_nara.py
python3 nara/check_nara.py --watchShows process, completed stages, latest merged harvest record count, disk usage, and recent log lines.
Smithsonian delivers files to s3://dpla-hub-si bimonthly (Feb, Apr, Jun, Aug, Oct, Dec). It requires a preprocessing step (fix-si.sh) before the standard pipeline and has optional human review checkpoints at each stage.
python3 smithsonian/launch_smithsonian.pyStages run in order with a confirmation prompt between each:
- Download — syncs from
s3://dpla-hub-sito EC2 - Preprocess — runs
fix-si.shto normalize raw files (checkpoint) - Harvest — runs
harvest.sh smithsonian(checkpoint) - Mapping — runs
ingest.sh smithsonian --mapping-only(checkpoint) - Pipeline — runs
ingest.sh smithsonian --resume-from enrichment
python3 smithsonian/launch_smithsonian.py --auto # skip checkpoints
python3 smithsonian/launch_smithsonian.py --date 2026-04-03 # specific delivery date
python3 smithsonian/launch_smithsonian.py --start-at mapping # resume from a stageMonitor with:
python3 smithsonian/check_status_smithsonian.py
python3 smithsonian/check_status_smithsonian.py --watchInternet Archive delivers a SQLite .db file directly (not via S3). File recieved to tech@dp.la and downloaded to local machine.
python3 community-webs/launch_cw.py --db ~/Downloads/community-webs.db --fullThis:
- Uploads the
.dbtos3://dpla-scratch/community-webs/ - Downloads it to EC2 at a temp path
- Runs
community-webs-ingest.sh --db=<path> --fullon EC2: export → harvest → mapping → enrichment → jsonl → S3
Without --full, only export and harvest run. Re-run with --full and --skip-export to continue:
python3 community-webs/launch_cw.py --full --skip-exportSkip flags for resuming after a failure:
--skip-upload # .db already in s3://dpla-scratch/community-webs/
--skip-export # ZIP already on EC2, skip straight to harvestMonitor with:
python3 community-webs/check_cw.py
python3 community-webs/check_cw.py --watchScripts send notifications to #tech-alerts via the ingest EC2. Configure webhooks in .env:
SLACK_WEBHOOK= # #tech-alerts
SLACK_TECH_WEBHOOK= # #tech (hub-complete with @here)
SLACK_ALERT_USER_ID= # your Slack member ID for @mention on failures
Clock skew / InvalidSignatureException from AWS CLI
AWS rejects requests when your system clock is off. Fix with:
sudo sntp -sS time.apple.comThen re-run or use --resume / --cluster-id to pick up where you left off.
Ingest failed mid-run
Use --resume-from <stage> in launch_ingest.py to restart from mapping, enrichment, or jsonl without re-harvesting.
Cluster monitoring crashes
Your IAM user needs EMR read permissions. Resume with --cluster-id j-XXXXX once permissions are sorted.
Slack notifications not sending
Verify ~/.dpla-secrets.env exists on the ingest EC2 with SLACK_BOT_TOKEN and SLACK_CHANNEL set.
sbt assembly fails
Make sure sbt is installed (brew install sbt on Mac) and batch-process-dpla-index is cloned at the path saved in your .env.
NARA credentials expired
The [nara] profile in ~/.aws/credentials on EC2 uses time-limited keys. Request new ones from NARA and update the profile.