HPC User Policy
Research Cores Program HPC Access Policy
This policy establishes the framework for access to high-performance computing (HPC) resources managed by the Research Cores Program (RCP) within the Office of Research, Creativity, and Economic Development (RCED) at New Mexico State University (NMSU).
Unlike NMSU’s previous Discovery cluster, which is managed by Information Technologies (IT), a new GPU cluster will be managed directly by the RCP. This shift ensures that the system is built, operated, and expanded in close alignment with faculty research needs, with RCP serving as the conduit between faculty researchers and IT technical staff.
Shared, Researcher-Led Ecosystem
RCP provides an environment where faculty can confidently contribute compute nodes, GPUs, software, storage, and other capacity to a shared ecosystem. Key principles:
- Universal Access: All NMSU researchers and students will have free access to HPC compute and storage.
- Faculty-driven growth: Researchers build and shape the system to serve evolving needs.
- Guaranteed priority for contributors: Contributors receive priority access to their resources.
- Robust support: Technical services are provided in partnership with IT’s HPC team.
Condominium Model for Contributions
The cluster will operate under an investor-based condominium model:
- Faculty/PI contributors are “investors” and retain priority access to their nodes.
- Idle contributed nodes are available to non-owners, with the possibility of pausing those jobs if the owner submits work.
- Each hardware contribution includes a corresponding partition of storage for archiving. In particular, each contributed compute node includes 10 TB of complementary archival storage.
- Each condominium model contribution will include a sustainability fee of 10% of total node code, unless waived by RCP.
- Typical node lifespan is 5 years but may operate beyond lifespan when operational.
- The following stakeholders may participate in partial node purchases:
- Individual faculty members
- Multi‑PI research teams
- Departments, centers, and institutes
- External collaborators contributing through existing agreements
- In partnership with the RCP, IT will provide backend support. This model promotes:
- Sustainability through distributed investment and shared maintenance.
- Priority access for contributors while offering bursting capacity to the broader community.
- Grant integration, as faculty are encouraged to budget computing resources into external proposals, ensuring long-term viability.
High-Performance Computing (HPC) Priorities & Allowances
NMSU offers a tiered access system for its HPC, prioritizing equitable access to computational resources. While all have access to the HPC, priority access is determined by the level of contributions made to the condominium model as defined below.
Condominium Model (for priority access and data storage beyond free allowance):
For faculty requiring additional computational resources or in need of priority access, NMSU offers a modular access or condominium model for both computing and storage. These faculty will have priority access to the computational nodes they contribute, though the nodes will be available for general purpose use when idle. The storage will be reserved for only those contributing to the HPC. A 10% sustainability fee will be applied to any node purchase. This ensures longevity of the cluster and also provides contributors with sustained priority access on any replacement hardware purchased through sustainability funds.
- Tiered Queue Access: The HPC will be available through a tiered access system:
- Priority Access: Contributors to the condominium model, which can be individuals or research consortiums of faculty, staff, and students, receive priority access. The contributors of additional compute have priority access for the specific nodes purchased as an individual or consortium. Contributors also have access to use other nodes in the cluster when idle in addition to their own contributed nodes.
- Community Tier Access: NMSU faculty, staff, and students not contributing nodes have access to the HPC but do not have priority access to any node or set of nodes.
- Collaborator Access: A mechanism is already in place for collaborators of NMSU faculty to access the HPC, with collaborators of contributors to the cluster receiving priority access. The complete process for accessing the HPC is outlined on the HPC website.
- Computational Allowances:
- For contributors to the condominium model: Faculty, staff, and student members are eligible for an unlimited and instant allowance of compute time and 10 TB storage. There are no direct costs for contributors, who submit an application via a managed web form, with assistance from the NMSU Director of Research Computing and Data Science for access. The priority access extends to the lab members or team of a contributor to the cluster. However, to retain priority access, the managing faculty or staff member must submit an access request.
- For NMSU Faculty, Staff, and Students: Non-contributing NMSU faculty, staff, and students are eligible for an unlimited allowance of compute time and 100 GB of storage, though with lower priority access compared to No direct costs apply to NMSU community members. NMSU faculty or staff submit an application via a managed web form, and the NMSU Director of Research Computing and Data Science assists with access.
With increasing costs associated with GPU‑based computing hardware, the Research Cores Program (RCP) will support fractional financial contributions toward the purchase of shared compute nodes. This allows researchers or research groups with smaller available funding to become contributors within the condominium model and receive proportional priority access. This expanded model ensures equitable participation, sustainable system growth, and broader buy‑in across disciplines. Participants receive contributor status proportional to their share and may combine funds from multiple indexes.
Structure of Fractional Ownership
A compute node may be divided into contribution shares, each representing a percentage of the hardware purchase price.
- A shared node may be divided into any reasonable number of shares (e.g., 2 50% shares or 40/60% share, four 25% shares, ten 10% shares, etc.).
- Minimum share size is 10% of the total hardware cost unless otherwise approved by RCP.
- Fractional contributors receive storage in proportion to their investment. Additional archival storage may be purchased at standard RCP storage rates.
- RCP manages all purchasing, warranty, and technical review requirements.
- The condominium model’s sustainability fee (10%) applies proportionally to each share, just as with full node purchases with sole ownership.
- RCP remains sole administrator of hardware, configuration, and lifecycle management to ensure that shared nodes remain compatible with cluster-wide standards.
Priority Access for Fractional Contributors
Fractional contributors receive priority access based on their ownership percentage.
- Priority scheduling is proportional to the contributor’s share of the node.
- When multiple fractional owners submit jobs concurrently, scheduling honors share proportions (e.g., a 40% owner receives up to 40% of the node’s dedicated queue capacity).
- When all fractional contributors are not using their priority share, idle cycles become available for the general community tier.
- Community-tier jobs on shared nodes will be paused if any fractional owner submits new work.
This system ensures benefits even for low‑dollar investors while preserving fairness among larger contributors.
Combining Small Contributions into Shared Nodes
For researchers with limited funding, RCP will maintain a Shared Node Investment Pool, enabling participants to contribute funds toward pooled nodes that RCP constructs once threshold funding is met.
- RCP announces node specifications and estimated costs annually.
- Researchers contribute any amount ≥$5,000.
- Once funding reaches the purchase threshold, RCP procures the node on behalf of all contributors.
- Priority access is calculated as:
Contributor Priority = (Individual Contribution) / (Node Cost) - If a contributors purchases a share of a node and, at a later time makes additional contributions to the HPC, the priority access is adjusted each time a new contribution is registered.
- To maintain clarity for contributors, RCP will provide a record of participating investors and their ownership percentages
This model supports widespread participation regardless of funding level or access at one point in time.
Benefits of Allowances:
This faculty allowance can be integrated into start-up and retention letters for new and existing faculty. It can also serve as an institutional cost-share on grant proposals, demonstrating strong university support for research. Consulting staff will be available at no charge to help faculty learn about the facilities and optimize their research software. These allowances apply for contributors of full or fractional nodes.
- The only direct cost to faculty members in this model is the price of the computing servers or storage "shelves" they wish to add to the NMSU cluster.
- The RCP covers other significant one-time and recurring costs, including system administration staff, storage, network infrastructure, and consulting.
- Researchers are provided compute resources based on their contribution and gain access to burst capacity beyond their purchased contribution. A portion (10%) of the charges for the purchased contributions will sustain the management, warranty renewals, and upgrades of the system as expansions are made. This ensures sustainability and access to these contributed compute resources.
- The purchased contributions will be managed by the HPC team of IT, thus relieving faculty of maintenance and administrative duties.
- Storage Structure and Resource Tiers: Faculty can contribute full or fractional nodes to the computing cluster to gain additional compute resources and storage.
Table 1: Storage levels and allocation methods
|
Storage Tier |
Description |
Allocation Method |
|
Home directory |
Small size storage for user accounts, code, and critical files. |
Standard quota for all users, regardless of node contribution. |
|
Project storage (100 Gb) |
Shared space for a research group to store application software and data. |
A base amount of project storage is given to all HPC users, with additional capacity purchased or granted in proportion to the number of nodes contributed. For groups with larger data needs, additional storage is available for purchase. |
|
Archival storage (10 Tb per contribution) |
Long-term, low-cost storage for data not actively being used but needs to be retained. |
Access and pricing are often negotiated on a project-by-project basis, with institutional subsidies sometimes available. |
- Faculty agree to share unused computing capacity with other users when it is not actively being utilized.
- Benefits for Recruitment, Retention, and Grant Proposals:
- Deans or Chairs can commit funds for faculty to invest in the HPC and storage, with these commitments explicitly noted in start-up or retention letters.
- The university's support for the HPC, particularly its coverage of recurring costs, can be leveraged as institutional cost-share on grant proposals, which federal funding agencies have indicated is important.
- For start-up letters, a statement should indicate an allocation (e.g., $50,000) for purchasing multiple high-performance computing nodes and large-scale storage, explicitly mentioning the campus's coverage of ongoing system administration, high-speed network access, parallel storage, and consulting.
- For grant proposals, this support can be detailed as significant yearly campus contributions covering system administration, professional HPC-IT staff, and data center housing costs (collocation, power, cooling, network connections).
Institutional Impact
The new GPU cluster expands NMSU’s computational capabilities and solidifies its role as a leader in AI, ML, and data-driven research. Through institutional support, collaborative governance, and diversified funding, this system will:
- Enhance research capacity across disciplines.
- Strengthen partnerships with academia, government, and industry.
Support training and workforce development in computational research. In the following, additional details about this HPC facility, computing allowances, access policy, management of resources, and key service areas are outlined below.
Key Service Areas of the HPC as administered by the Research Cores Office in partnership with NMSU IT
- High-Performance Computing (HPC) Cluster:
- Provide high-performance computing (HPC) resources, featuring high-speed, low-latency interconnect and a high-speed parallel filesystem.
- Ensure free access through the Faculty/Staff/Student Computing Allowance program.
- Offer infrastructure and administration costs for modular contributions, with unused capacity available to the broader campus community.
- Allow modular contributors to "burst" beyond their purchased hardware capacity.
- Provide a low-cost large-scale data storage service through the condominium access model.
- Supports reference and learning modules.
- Managed in partnership with NMSU IT
- User Experience & Engagement:
This section outlines strategies to lower the barrier to entry and enhance accessibility for a diverse user base, aiming to provide equitable access to computational resources. With support for lines of graduate assistants (GA), the user experience and training opportunities could grow and be sustainable.
- User Interface & Usability:
- Job Request System: Implement a dual-interface system, including a simplified web-based portal with guided workflows for new users and a traditional command-line interface (CLI) for advanced users. The goal is an intuitive entry point for users with less technical experience while retaining full functionality for experts, with the new web portal prototyped by Q1 2026.
- Documentation: All user guides and tutorials will be rewritten using plain language and clear examples. Core documents are targeted for completion by the end of 2026.
- Training & Workshops:
- Tiered Training Program: A structured training program with three tiers will be developed based on the already available trainings. Creating a lower level “entry” set of modules would help those new to HPC environments and a higher level set of tutorials would complement the existing trainings and update with any new use cases. This ensures users can find training matching their current skill level, with a program launch planned for 2026.
- On-Demand Resources: A library of short, video-based tutorials and webinars on key topics (e.g., submitting a job, using specific software packages, debugging common errors) will be created. The goal is to provide flexible, on-demand learning resources for self-paced instruction, with the first series of videos available by mid-2026.
- User Support:
- Multi-Channel Support: Support will be offered through multiple channels, including an easy-to-use ticketing system, livechat office hours for immediate help, and dedicated email support, ensuring users can get help comfortably. The support will provide mechanisms for users to reach out to the team out of the Research Cores Program and share resources, ideas, and information with fellow users. This will develop a community around HPC users at NMSU and other affiliates for support. All channels will be clearly advertised on the website by mid-2026.
- Dedicated Outreach: HPC specialists will be assigned as liaisons to specific university departments and/or colleges to provide tailored support and promote HPC resource use. This liaison program aims to build relationships and make HPC resources less intimidating to new user groups, established by late 2026.
- Monitoring & Evaluation:
- Biannual user surveys will gather feedback on the accessibility of services, documentation, and support.
- Key metrics, such as the number of new user accounts, attendance at training sessions, and average time to resolve support tickets, will be tracked.
- An annual report on progress in improving user accessibility will be published, and this plan will be reviewed and updated annually based on user feedback and changing needs. This will be completed by the Director of Research Computing and Data Science.
Data Storage Policy
Only HPC staff and HPC users can store data on HPC systems. This includes both compute and archival storage.
Research data of the following types can be stored on the HPC drives:
- Input data for running and output data produced on HPC
- Executable files run on HPC and associated files, such as, source code, object files, user libraries, include files, make files
Acceptable Use Policy
The NMSU HPC system is a specialized resource intended for university research and instruction, and users must adhere to specific rules to ensure security, fairness, and legal compliance. The faculty member directing a project is ultimately responsible for ensuring all associated accounts use the services appropriately. Failure to comply with these policies can lead to disciplinary actions, including account suspension, termination, and potential legal action. The HPC support team monitors system usage to ensure compliance with these policies. By using the HPC resources, you agree to abide by these rules. The following outlines key responsibilities for using the HPC.
- Account Security and Access
User accounts and login credentials are for individual use only and must never be shared, even with colleagues on the same project.
- You are responsible for all activities that occur under your account.
- Any security weaknesses or unauthorized use of your account must be reported immediately to the HPC support team or Help Desk.
- Users must not attempt to access data or systems for which they do not have authorization. This includes trying to bypass security measures or running programs designed to find security weaknesses, like password crackers.
- Proper Use of Resources
The systems are designated for university research and instruction. Commercial activities, illegal activities, personal projects, and cryptocurrency mining are strictly prohibited.
- Users should not run jobs on login nodes, as these are reserved for tasks like editing files and submitting jobs. Processes that consume significant memory or CPU on login nodes will be terminated without notice.
- Jobs should be optimized to use resources efficiently, and resource requests should be as accurate as possible.
- Data Management and Security
- Storing classified, sensitive, or controlled data (such as HIPAA, CUI, FERPA) is generally forbidden on standard HPC systems.
- Specialized, secure systems must be used for data regulated by HIPAA or NIH genomic data policies, and users must consult with administrators before handling such information. Controlled Unclassified Information (CUI) cannot be stored or processed on general research computing resources.
- You are responsible for protecting information in your account and are encouraged to maintain your own backups. It's important to note that files on scratch systems are not backed up and may be deleted automatically.
- Software and Legal Compliance
- Users must have proper authorization to use copyrighted files and must comply with all software licensing agreements.
- Unauthorized copying, distribution, or use of software, including cracked or hacked programs, is prohibited.
- Users must not engage in any activity that violates U.S. export control laws or other local, state, and federal regulations.