Quick Information
Supercomputing systems present complex challenges to personnel who design, deploy and maintain these systems. Standing up these systems and keeping them running require novel solutions that are unique to high performance computing. The success of any supercomputing center depends on stable and reliable systems, and HPC Systems Professionals are crucial to that success.
The Eighth Annual HPC Systems Professionals Workshop will bring together systems administrators, systems architects, and systems analysts in order to share best practices, discuss cutting-edge technologies, and advance the state-of-the-practice for HPC systems. This CFP requests that participants submit either papers, slide presentations, or 5-minute Lightning Talk proposals. Additionally reproducible artifacts (code segments, test suites, configuration management templates) which can be made available to the community for use are welcome for submissions either as a standalone submission or in addition to any paper or talk submissions.
Call for Participation
Paper submissions are open. Check out our Call for Participation. Please send our CfP to any interested parties. Submissions are via the SC26 submission site.
Social Event
TBD
Schedule
All times in Eastern Time
| Start | End | Description |
|---|---|---|
| 8:30 AM | 8:45 AM | Opening Remarks |
| 8:45 AM | 9:10 AM | Enabling Performance Isolation for Co-scheduled Workflows, Daniel Margala (LBNL NERSC), Jonathan Skone (LBNL NERSC), Christopher Samuel (LBNL NERSC) |
| 9:10 AM | 9:35 AM | Operationalizing Workflow Support on Slurm: A Site Toolkit for Choosing Nextflow Deployment Strategie, Nil Tianchen Mu (Arizona State University) |
| 9:35 AM | 10:00 AM | Obleth: Fair Multi-Tenant LLM Serving for Self-Hosted Campus Infrastructure, Jonathan Lee (Arizona Ste University) |
| 10:00 AM | 10:30 AM | Morning Coffee Break |
| 10:30 AM | 11:00 AM | HPC Infrastructure Management at Scale, Jeremiah Hoyle (Accenture), Nishit Patel (Shell), Daniel Rix (Accenture), Cory Kim (Accenture), Carl Dukatz (Accenture), Chris Young (Accenture), Michael Gujral (Shell), Adam Hough (Microsoft Corporation), Fokko Masselink (Shell), John Thiesfeld (Shell), Ronald Cogswell (Shell) |
| 11:00 AM | 11:30 AM | Ops Knowledge That Stops Evaporating: Building an Evergreen Documentation, Search, and Monitoring System on an Agent CLI, Adam Hough (Microsoft Corporation) |
| 11:30 AM | 12:00 AM | Subsurface to Stellar: HPC Enablement with Open OnDemand and ColdFront at bp and NASA, (Panel) |
| 12:00 PM | 1:00 PM | Lunch Break |
| 1:00 PM | 1:30 PM | Architecting Multi-Petabyte Data Migration in a Live HPC Environment: From Strategy to Execution, Deepa Phanish (Georgia Institute of Technology), Aiden Lambert (Georgia Institute of Technology), Aaron Jezghani (Georgia Institute of Technology), Ruben Lara (Georgia Institute of Technology), Rachel Lombardi (Georgia Institute of Technology), Kenneth Suda (Georgia Institute of Technology), Grigori Yourganov(Georgia Institute of Technology) |
| 1:30 PM | 2:00 PM | Experiences with and Best Practices around Cooling Distribution Units (CDUs), Jenett Tillotson (NSF NCAR) |
| 2:00 PM | 2:30 PM | The Evolution of the Coaraci HPC Infrastructure: Operational Lessons from a Brazilian Research Computing Center, Mateus Francisco (Unicamp) |
| 2:30 PM | 3:00 PM | Lunch Break |
| 3:00 PM | 3:20 PM | Faster Post-Outage Testing at Full Machine Scale, Rory Kelly (NSF NCAR), Ben Matthews (NSF NCAR), Jenett Tillotson (NSF NCAR) |
| 3:20 PM | 3:40 PM | Security and Day-2 Resilience as First Class Control Plane Primitives for Heterogeneous Supercomputers, Sadaf R. Alam (University of Bristol), Alex Joseph Lovell-Troy (Los Alamos National Laboratory), Mark Klein (Swiss National Supercomputing Centre) |
| 3:40 PM | 4:00 PM | Death by a Thousand Spreadsheets: Consolidating Infrastructure Data with NetBox, Jeffrey Winters (Neuvys Technology), Rebecca Pinheiro (Neuvys Technology) |
| 4:00 PM | 4:20 PM | Teaching HPC to Advanced High School Students at a Summer Program, Mike Renfro (Tennessee Tech University) |
| 4:20 PM | 4:30 PM | Chapter Updates and Closing Remarks, TBD |
Topics of Interest
Here are some topics of interest for this group. Note that these are here to indicate direction, not to disallow other related topics.
- Cluster, configuration, or software management
- Cybersecurity and data protection
- Performance tuning/Benchmarking
- Resource manager and job scheduler configuration
- Monitoring/Mean-time-to-failure/ROI/Resource utilization
- HPC storage solutions
- High speed/ Low Latency networking
- Composable infrastructure and containers
- Elastic workloads or optimizations for workload types
- Web-based cluster front ends
- Challenges with AI workloads (GPU management, Interconnect, Data Movement)
Example paper ideas might be:
- Best practices for job scheduler configuration
- Advantages of cluster automation
- Managing software on HPC clusters
Calendar
| Event | Date |
|---|
Submissions
You can use the following link to submit your presentation for review: SC26 Submission Site
Organizing Committee
| Position | Name | Affiliation |
|---|---|---|
| Workshop Chair | Michael Hartman | Stanford |
| Workshop Co-Chair | Jason Blair | P&G |
| Program Chair | David Clifton | Ansys |
| Organizing Committee | ||
| Blaise Hartman | NASA | |
| Betsy Hillery | Purdue University | |
| Hon Wai Leong | DDN | |
| John Legato | NIH | |
| Gary Skouson | Penn State University | |
| Kurt Maier | PNNL | |
| Stephen Fralich | Boeing |
Program Committee
| Name | Affiliation |
|---|
Publication Information
All accepted papers and artifacts will be published on GitHub and archived with a DOI in Zenodo. You can view the previous years presentations here HPCSYSPROS SC23 Workshop Proceedings
Contact Information
If you need to contact us, send email to SIGHPC SYSPROS.
Links
- SC HPC Syspros Mailing List - you should join!
- Join our SIGHPC SYSPROS Slack team
- Email us with any questions
