HN user

jfkfif

28 karma
Posts0
Comments20
View on HN
No posts found.

It is not a reasonable assumption to compare your local cluster to the largest clusters within DOE or their equivalents in Europe/Japan. These machines regularly run at >90% utilization and you will not be given an allocation if you can’t prove that you’ll actually use the machine.

I do see the phenomenon you describe on smaller university clusters, but these are not power users who know how to leverage HPC to the highest capacity. People in DOE spend their careers working to use as much as these machines as efficiently as possible.

HPC, or scientific computing more generally, are slow when it comes to adopting modern software development practices. Most researchers used to only care about results above all, rather than reliability or reproducibility for other users. DOE initiatives [0] in the past couple years have done a lot to change this, though there are significant hurdles from a security and logistical standpoint [1]. For instance, to uncover issues in a PR, you might need to run the code at >64 nodes on a shared system.

[0] https://ecp-ci.gitlab.io/docs/admin/jacamar/introduction.htm...

[1] https://arxiv.org/abs/2303.17034

Fortran 2023 3 years ago

LLNL has ported most if not all of their codes to C++, but I think LANL is still full on Fortran

Chapel 1.32 3 years ago

please email us at hpccenter@stanford.edu, and ask to be put in touch with someone at PSAAP. I would be happy to help!