One time I had a ticket where the user SWORE WITH THEIR WHOLE HEART nothing changed between job 1 and job 2 that failed.
Nothing.
Until finally … “well of course it’s a new input!”
Me: How different is this new input?
Oh it’s MUCH bigger.
🤦♀️🤦♀️🤦♀️
I miss the "job arrays seem complicated so I just scripted a loop to sbatch 80,000 individual jobs ..." conversations. T
Today most of my requests are due to #aws#parallelcluster auto-scaling failing for quota or "insufficient ec2 capacity" errors -- both error modes not easily visible to users