Back to Search Start Over

Best practices for management and operation of large HPC installations.

Authors :
Lathrop, Scott
Mendes, Celso
Enos, Jeremy
Bode, Brett
Bauer, Gregory
Sisneros, Roberto
Kramer, William
Source :
Concurrency & Computation: Practice & Experience; 8/25/2019, Vol. 31 Issue 16, pN.PAG-N.PAG, 1p
Publication Year :
2019

Abstract

Summary: To achieve their mission and goals, HPC centers continually strive to improve the effectiveness of their resources and services to best serve their constituencies. Collectively, the community has learned a great deal about how to manage and operate HPC centers, provide robust and effective services, and develop new communities as well as about other important aspects. Yet, cataloguing best practices to help inform and guide the broader HPC community is not often done. To improve the situation, the Blue Waters project has documented sets of best practices that have been adopted for the deployment and operation over the past five years of the Blue Waters leadership system, a large Cray XE6/XK7 supercomputer at NCSA. Those practices, described in this paper, cover aspects of managing and operating the system and its resources, supporting its users, and expanding the diversity of applications and communities. Although the technical practices are sometimes discussed relative to Cray systems and leadership‐scale systems, we believe that they would benefit the deployment and operation of other large HPC installations as well. [ABSTRACT FROM AUTHOR]

Details

Language :
English
ISSN :
15320626
Volume :
31
Issue :
16
Database :
Complementary Index
Journal :
Concurrency & Computation: Practice & Experience
Publication Type :
Academic Journal
Accession number :
137639859
Full Text :
https://doi.org/10.1002/cpe.5069