Environment
EGAPx version: 0.5.2
Executor: SLURM
Genome: ~2.5 Gb mammal
RNA-seq: 9 paired-end Illumina samples, ~7.6 GB gzipped per file
I'm running EGAPx on an HPC with the default slurm.config template causes star_wnode to write ~5.5 TB of temporary files to the Nextflow work directory for a single RNA-seq sample, filling shared storage.
The cause is that the slurm.config template sets TMPDIR, TMP, and TEMP to $PWD/tmp, but star_wnode defaults to /dev/shm for its work area (confirmed from star_wnode --help). Setting these env vars overrides that and sends all temp output to disk instead of RAM.
Our /dev/shm has 504 GB available, plenty for this purpose.
Configs
slurm.config:
singularity {
enabled = true
autoMounts = true
cacheDir = "$PWD/singularity"
envWhitelist='https_proxy,http_proxy,ftp_proxy,DISPLAY,SLURM_JOB_ID,SINGULARITY_BINDPATH'
}
env {
SINGULARITY_CACHEDIR="$PWD/singularity"
SINGULARITY_TMPDIR="$PWD/tmp"
DEBUG_STACK_TRACE_LEVEL="Warning"
EXCEPTION_STACK_TRACE_LEVEL="Warning"
DIAG_POST_LEVEL="Trace"
DEBUG_CATCH_UNHANDLED_EXCEPTIONS="0"
NXF_TEMP = "$PWD/tmp"
// TMPDIR, TMP, TEMP removed — these were overriding star_wnode's
// default work area from /dev/shm to disk, causing 5.5TB disk usage
// CONN_DISPD_DISABLE="1"
// CONN_LBSMD_DISABLE="1"
}
process {
executor = 'slurm'
queue = 'long'
queueSize = 20
clusterOptions = ' --ntasks=1 '
memory = '32 GB'
cpus = 8
time = '72h'
ext.threads = 8
ext.split_jobs = 20
withLabel: 'small_mem' {
memory = '8 GB'
}
withLabel: 'med_mem' {
memory = '64 GB'
}
withLabel: 'large_mem' {
memory = '128 GB'
}
withName: 'run_star' {
memory = '256 GB'
maxForks = 1
cpus = 20
ext.threads = 20
}
withName: 'fetch_ortholog_references' {
memory = '32 GB'
}
}
example.yaml:
yamlgenome: /path/to/GCA_028749985.3_mPumCon1.1.hap1_genomic.fna
taxid: 9696
short_reads: /path/to/short_reads.txt
tasks:
star_wnode:
star_wnode: "-preserve-star-logs -workers 1 -cpus-per-worker 20"
Environment
EGAPx version: 0.5.2
Executor: SLURM
Genome: ~2.5 Gb mammal
RNA-seq: 9 paired-end Illumina samples, ~7.6 GB gzipped per file
I'm running EGAPx on an HPC with the default slurm.config template causes star_wnode to write ~5.5 TB of temporary files to the Nextflow work directory for a single RNA-seq sample, filling shared storage.
The cause is that the slurm.config template sets TMPDIR, TMP, and TEMP to $PWD/tmp, but star_wnode defaults to /dev/shm for its work area (confirmed from star_wnode --help). Setting these env vars overrides that and sends all temp output to disk instead of RAM.
Our /dev/shm has 504 GB available, plenty for this purpose.
Configs
slurm.config:
singularity {
enabled = true
autoMounts = true
cacheDir = "$PWD/singularity"
envWhitelist='https_proxy,http_proxy,ftp_proxy,DISPLAY,SLURM_JOB_ID,SINGULARITY_BINDPATH'
}
env {
SINGULARITY_CACHEDIR="$PWD/singularity"
SINGULARITY_TMPDIR="$PWD/tmp"
DEBUG_STACK_TRACE_LEVEL="Warning"
EXCEPTION_STACK_TRACE_LEVEL="Warning"
DIAG_POST_LEVEL="Trace"
DEBUG_CATCH_UNHANDLED_EXCEPTIONS="0"
NXF_TEMP = "$PWD/tmp"
// TMPDIR, TMP, TEMP removed — these were overriding star_wnode's
// default work area from /dev/shm to disk, causing 5.5TB disk usage
// CONN_DISPD_DISABLE="1"
// CONN_LBSMD_DISABLE="1"
}
process {
executor = 'slurm'
queue = 'long'
queueSize = 20
clusterOptions = ' --ntasks=1 '
memory = '32 GB'
cpus = 8
time = '72h'
ext.threads = 8
ext.split_jobs = 20
}
example.yaml:
yamlgenome: /path/to/GCA_028749985.3_mPumCon1.1.hap1_genomic.fna
taxid: 9696
short_reads: /path/to/short_reads.txt
tasks:
star_wnode:
star_wnode: "-preserve-star-logs -workers 1 -cpus-per-worker 20"