Submitting more applications increases your chances of landing a job.
Here’s how busy the average job seeker was last month:
Opportunities viewed
Applications submitted
Keep exploring and applying to maximize your chances!
Looking for employers with a proven track record of hiring women?
Click here to explore opportunities now!You are invited to participate in a survey designed to help researchers understand how best to match workers to the types of jobs they are searching for
Would You Be Likely to Participate?
If selected, we will contact you via email with further instructions and details about your participation.
You will receive a $7 payout for answering the survey.
Job Purpose
Provide high?throughput, consistent storage tiers (Scratch/HPS + Object) for large?scale training data ingest, checkpoints, and inference artifacts.
Roles & Responsibilities
Design/expand Lustre/BeeGFS HPS; NVMe?oF and Object (S3) tiers; align with AI dataflow and GDS.
Establish namespace, OST/MDT layout, stripe/RAID policies; tiering for warm/cold datasets.
Capacity/performance planning; rebalance and failover testing; automate snapshots and checkpoint retention.
Proactive detection of hot spots and metadata contention; schema for small?file handling.
Tune RDMA paths, page cache, IO schedulers; validate end?to?end I/O profiles for LLM training/inference.
Lead P0/P1 critical incident response for enterprise-scale AI/ML storage infrastructure supporting NVIDIA GPU clusters, ensuring rapid service restoration and minimal impact to business-critical workloads.
Act as the storage SME during major incidents involving BeeGFS, Lustre, GPFS (IBM Spectrum Scale), NFS, NVMe-oF, Parallel File Systems, and Object Storage platforms.
Perform deep-dive troubleshooting and resolution of storage performance degradation, metadata bottlenecks, I/O latency spikes, filesystem corruption, capacity exhaustion, and hardware failures.
You'll no longer be considered for this role and your application will be removed from the employer's inbox.