policy – FASRC DOCS https://docs.rc.fas.harvard.edu Tue, 28 Jul 2026 18:13:03 +0000 en-US hourly 1 https://wordpress.org/?v=7.1 https://docs.rc.fas.harvard.edu/wp-content/uploads/2018/08/fasrc_64x64.png policy – FASRC DOCS https://docs.rc.fas.harvard.edu 32 32 172380571 Support Expectations and Hours – Policy https://docs.rc.fas.harvard.edu/kb/support-and-hours-policy/ Tue, 28 Jul 2026 17:56:50 +0000 https://docs.rc.fas.harvard.edu/?post_type=epkb_post_type_1&p=30285 FASRC support is available to users of our services Monday through Friday 9am-5pm.

To submit a ticket, visit our Contact page. Someone will get back to you as soon as possible.

Please note that FASRC is not staffed on nights, weekends, holidays, and over the winter break but maintains monitoring to notify us of critical issues out of hours

]]>
30285
FASRC AI Facilitation https://docs.rc.fas.harvard.edu/kb/fasrc-ai-facilitation/ Wed, 22 Jul 2026 06:51:38 +0000 https://docs.rc.fas.harvard.edu/?post_type=epkb_post_type_1&p=30249 FASRC Policy and Safety Guidelines for AI Workflows

FASRC allows the use of AI tools on the Cannon and FASSE FASRC clusters. Our users are free to install AI tools they need to execute their AI-based workflow under their cluster profile and provide it with data. However, we urge our users to be aware of the guidelines that the University has put forward for the use of such tools at Harvard. We do ask that for all HUIT-supported LLMs, our users register apps and access API keys that are under Harvard agreement prior to accessing those LLMs on the cluster.  

Additionally, if you are using AI agents for your work, you need to know how to use them responsibly on the cluster without compromising your data security or your user profile on the cluster. The AI Agents on the FASRC Clusters policy page walks you through the main points you need to be aware of in order to use AI agents responsibly on the cluster while maintaining the data integrity and data privacy of your work.

Note: Sometimes Chrome does not work for HUIT websites, especially if you are *not* on the University VPN. In that case, you can access those websites using Firefox.

Best Practices for AI workflows on FASRC Clusters

In order to execute an AI/ML workflow successfully on the cluster, it is important to understand:

  • How to safely install and launch a tool/software on the cluster that’s needed for your workflow
  • How to request resources for it properly
  • What are some of the pitfalls to be mindful of

This section provides a landing area for all the relevant documentation that will help you implement your AI/ML workflow on the cluster correctly. Before executing your AI workflow on the FASRC clusters, make sure to follow the safety guidelines mentioned above, and FASRC guidelines for launching an AI tool such as OpenAI’s ChatGPT.

How to install and utilize AI/ML tools or platforms? 

To successfully execute your AI/ML workflow on the cluster, you may need a variety of tools or platforms. The list below captures various tools and platforms that are currently available on FASRC clusters. Based on your needs, you can go through the relevant documentation to get guidance on how to install and execute these tools on the cluster, or utilize these platforms for your work.

* Anthropic

* HeavyAI

* Knime

* OpenAI

* Python Package Installation

* PyTorch

* Tensorflow

* VSCode

* Cursor

* AI Extensions

How to request resources for AI workflows?

This will typically require you to request GPUs and utilize them efficiently. GPU Computing on FASRC Clusters walks you through:

  • The resource allocation process for GPUs for a batch and an interactive job,
  • How to load and work with CUDA modules,
  • Example codes for GPU computing on the cluster
  • How to monitor your job’s performance on a GPU node.

In addition to that, one must be mindful of their job’s efficiency. A detailed description of what it means to run a job efficiently on the cluster is provided on Job Efficiency and Optimization Best Practices.    

Lookout For:

In addition to following the best practices mentioned above, one must be aware of some of the pitfalls associated with executing AI workflows on the cluster. For example:

  1. Are you using an AI tool or an extension via a code editor, such as VSCode or Cursor, on the login node?
  2. Have you inadvertently given permission to the AI tool to delete files or folders?
  3. Are you aware of prompt injection attacks while using an LLM-based application?
  4. What steps have you taken to secure your data against data leakage?
  5. Are you aware of your data security level?

Training

We regularly hold training sessions to support AI-based workflows on the cluster.  Please refer to our Training Calendar for upcoming sessions, and Training Material to access AI/ML computing documentation and presentation slide decks &/or video recordings of our previous training sessions.

]]>
30249
AI Agents on the FASRC Clusters https://docs.rc.fas.harvard.edu/kb/ai-agents/ Wed, 08 Jul 2026 13:56:12 +0000 https://docs.rc.fas.harvard.edu/?post_type=epkb_post_type_1&p=30201 Responsible Use of AI Agents
  • All users of the FASRC clusters and associated resources are bound by Harvard University guidelines on AI use.
  • Using personal accounts or API keys for work on the FASRC cluster is not in accordance with Harvard policy and guidelines.
  • Never run agents with elevated privileges.
  • You, the FASRC account holder, are responsible for the actions of your agents, These processes are running with your privileges.
  • Like you, your agent should behave responsibly on the cluster following customer and responsibilities and acceptable use guidelines as well as any data use requirements places on you, your lab, your PI, or the data being accessed by an agent.
  • You, the FASRC account holder, are responsible for your agent’s access to data that is in any way restricted and for its actions on any file systems resulting in data loss or data leak.
  • Limit access. Run agents in a sandbox or container where possible. If not possible, consider whether this action increases the chances of unintended data loss, leaks, or other un-wanted outcomes.
  • You, the FASRC account holder, are responsible for reviewing any AI and agentic skills before use.
  • You, the FASRC account holder, are responsible for monitoring your agents, their API usage, cluster usage, and any associated costs they may be incurring on behalf of you, your lab, or the university.
    • Monitor your agent sessions actively.
    • Limit the number of agents on login nodes.
    • Don’t run multiple agents without monitoring..
    • Terminate processes that behave unexpectedly.
    • If agents become disruptive, we may introduce automations to moderate their activities so that other users are not affected.

This document covers some, but not all, concerns and contingencies around the use of AI agents on the FASRC clusters.

If you have concerns or questions, please contact FASRC or consider coming to Office Hours.

Data Integrity and Data Privacy

You are responsible for the actions of your agents, therefore you are responsible for any breach of a DUA or other instrument which restricts access to data should that data be leaked, copied, removed, etc. by your agent. If in doubt, consult your PI when running agents in FASSE or on any data which has a data use agreement or other restrictions.

For questions about data use agreements you may be bound to, consult with your PI, OVPR or contact FASRC Research Data Management.

Related Links:

 

]]>
30201
Acceptable Use Policy https://docs.rc.fas.harvard.edu/kb/acceptable-use/ Wed, 28 Jan 2026 18:00:31 +0000 https://docs.rc.fas.harvard.edu/?post_type=epkb_post_type_1&p=29426 FAS Research Computing (FASRC) cluster access and usage is intended only for legitimate purposes which benefit research at Harvard University.  Access must be authorized by the faculty or management of the FAS or those of our partner schools, and by the staff of Research Computing.  Account access should only be granted for the purposes necessary to accomplish the goals of Harvard University and its research projects.  All active FAS RC account holders are subscribed to our notifications mailing list which is a requirement for all users.

Billing

Cluster usage and additional resources such as storage may be subject to charges to the PI, school, or department.  All billing is done exclusively via Harvard internal billing codes at the Tub/school level.

See our Data Storage Billing documentation.

Academic and Administrative Use

The FASRC clusters (Cannon and FASSE) are for research only and cannot be used for academic purposes. Harvard provides an Academic Cluster for those purposes.

FASRC cluster storage is for research data and results and is not suitable for administrative data storage.

Accounts

Accounts and account credential sharing is not allowed under university policies and reasonable precaution should be taken to keep your account credentials secure and private. No university staff will ever ask for your password.  Additionally, a user may have only one account at FASRC. All individual account holders, whether Harvard affiliates or outside collaborators agree to be held accountable by the Harvard University Electronic Access and Information Security polices: http://huit.harvard.edu/information-technology-policies. In addition, researchers should make themselves familiar with the university research policies maintained by the Provost’s Office.

All account holders agree to respect requests from support staff around how they use the system. The support staff may, as needed, impose whatever policies are required to ensure the system runs effectively for all users of the system.

Data Security

The Cannon cluster is for data rated as Level 2 or below. Level 3 data must be secured and processed on the FASSE cluster and storage. Level 4 or above data is not allowed on any FASRC cluster or storage.

Please also review the FASRC Cluster Storage Policy for guidelines an best practices around storage.

Customs and Responsibilities

In addition, the FASRC clusters and storage are shared resourcesm so please familiarize yourself with our Cluster Customs and Responsibilities.

Additional Policies

]]>
29426
FASRC Cluster Storage Policy https://docs.rc.fas.harvard.edu/kb/fasrc-cluster-storage-policy/ Tue, 13 Aug 2024 19:46:19 +0000 https://docs.rc.fas.harvard.edu/?post_type=epkb_post_type_1&p=27525 Cluster storage offered and maintained by FASRC should only be used for research taking place on FASRC clusters.

Examples of data that can be stored on FASRC storage are:

  • Datasets
  • Code
  • Scientific software
  • Research results

Examples of data that should not be stored on FASRC storage include:

  • Clerical or lab administrative data
  • Data related to personnel, grant proposals, business operations, or general lab management
  • Data with personally identifiable or financial information 

FASRC storage filesystems are only approved for Data Security Level 1 (DSL1) and  DSL2 research data on the Cannon cluster. DSL3 data must be stored in the approved FASSE cluster project. Research data containing information classified as DSL 4 must be stored on an appropriate storage solution that is approved for DSL4 sensitive data.*

*A limited number of DSL4 projects exist in their own isolated environments

If it comes to the attention of the FASRC Staff that non research related data is being stored on the FASRC systems, we will alert the lab’s PI.

To view alternative storage options for administrative data, please refer to the FASRC website.  Additional information is also provided on the Harvard Security website regarding Data Security levels.

]]>
27525
FASRC Data Ownership and Access Policy https://docs.rc.fas.harvard.edu/kb/data-ownership-and-access-policy/ Wed, 17 Jul 2024 13:33:15 +0000 https://docs.rc.fas.harvard.edu/?post_type=epkb_post_type_1&p=27362 Data stored in FASRC lab directories (/n/holylabs/<labname>_lab) and other lab shares on FASRC storage are owned and managed by PIs or group owners. If written approval is provided to FASRC by the PI or group owner, FASRC can modify the data permissions to allow ownership and access as the PI or group owners require.

  • Additional approval by the original owner of the data will not be required. The PI or group owner of the storage folder is responsible for its use.
  • Self-service mechanisms are already in place permitting labs and groups to modify folder permissions. This policy expands upon available tools, affirming FASRC’s right to alter folder permissions, when requested.

Additional notes

  • Written approval will be required from the PI or group owner for FASRC to modify access without the original owner’s consent.
    • Explicit written permission needs to derive from the PI or group owner, or an individual approved by the PI to assume similar responsibilities (i.e. a General Manager or Storage Manager). Please see “FASRC Roles and Responsibilities” for more details.
  • Lab folders on FASRC storage will initially be created as group writable, which allows for easier collaboration, data migrations across platforms, and data cleanup.
    • A lab or group may choose to modify folder permissions, but the default setup will be group writable.
  • If a lab folder is not group writable, FASRC can modify permissions to make a folder group writable, as requested by the PI or group owner or an individual approved by the PI or group owner to assume similar responsibilities (i.e. a General Manager or Storage Manager).

Associated University Policies

]]>
27362
Web Scraping Policy https://docs.rc.fas.harvard.edu/kb/web-scraping-policy/ Wed, 22 Apr 2020 11:59:46 +0000 https://docs.rc.fas.harvard.edu/?post_type=epkb_post_type_1&p=23311 Web scraping is a contentious issue within research. While it is true that fair use provides for many uses of data gleaned from the Internet, in general this is applied to human information gathering, not programmatic machine scraping. That distinction makes the act of brute-force scraping an issue separate from fair use.

You, as a representative of Harvard, are not just using the source’s data, but also their servers, bandwidth, etc. in a way the source may not approve. This can lead to IP blacklisting and even legal action. So please tread carefully as your actions could negatively affect others.

If in doubt or in need of more authoritative guidance, please contact the Harvard Office of the General Counsel or Office of the Vice Provost for Research

If you are scraping for the purpose of train a GAI model, contact the Harvard Office of the General Counsel or Office of the Vice Provost for Research

Please be aware that merely being involved in academic pursuits does not exempt you from the usage policies of social media and other Internet platforms like Facebook, Twitter, etc.

Sensitive Data

If the data you are acquiring is considered sensitive, confidential, or contains human data, you will need to have this data reviewed for compliance before placing it on the FASRC cluster. If in doubt, you should always err on the side of caution and contact the Office of the Vice Provost for Research

 

Scraping data for use on the FASRC Cluster

If your research requires you to scrape content from the web, please review the following guidelines and suggestions.

We highly discourage using the cluster itself to scrape data. Due to its size and ease of parallelization of processes, the cluster is easily weaponized and your actions could have consequences for other researchers. Please seek another avenue for data acquisition first.

You should contact FASRC before commencing any scraping activity using the FASRC cluster.

It is highly preferable that you do the scraping elsewhere and then bring the data to the FASRC cluster for processing. If the data is sensitive, confidential, contains human data, or it is unclear, then this is a requirement. See ‘Sensitive Data’ above.

Also, if you are scraping for the purpose of training a GAI/LLM model, you should respect that site’s policies on this practice (this may be posted on the site, contained in a robots.txt file, or explicitly stated in their ToS). Even if you are doing the scraping manually, you should consider yourself the same as a bot and, if a site excludes GAI/AI bots, this also applies to you. Merely being an academic does not exempt you from following the wishes of a site and/or its members; your exfiltrated data could end up in other models thereby nullifying the source’s right to exclusivity/ownership. Please contact the Harvard Office of the General Counsel or Office of the Vice Provost for Research for further guidance.

Source Permission

If you are in doubt or have questions, please contact the Harvard Office of the Vice Provost for Research

Data on the Internet should not be programmatically (or ‘brute-force’) scraped using FASRC computing resources, even for academic research purposes, unless FASRC has given permission to proceed using the cluster or some system tied to the cluster, and:

A) The source provides an API for this purpose and any requirements they impose have been met.

B) The source allows/does not prohibit scraping in their terms of service or other public notice.

C) The source is the United States government and the data in question was generated with public funds and is publicly available without encumbrance. Further, that the site not be scraped using brute-force means if an API is provided.

D) The source has given you explicit permission in writing or via a secondary document spelling out that permission.

E) The source does not exclude/forbid your use-case, such as GAI or LLM training.

Data cannot be programmatically scraped using FASRC computing resources if the source has explicitly forbidden scraping in their terms of service and written permission to do so cannot be obtained. In such a case, you should investigate other options for acquiring this or similar data.

Throttling and Blacklisting

Scraping content from websites using highly parallelized processes, even with unfettered permission from the source, should be avoided. Doing so runs the risk of having the cluster, or even the university’s, IP range blacklisted. This could have an undesirable effect on other network and cluster users. Please ensure your processes pull data at a reasonable rate unless you explicitly have written approval from the data source to download more aggressively and assurance that this will not lead to blacklisting from them or their upstream provider.

Related:

Harvard Office of the Vice Provost for Research

US Data.gov Data Harvesting Information

Archive.org Scraping

 

]]>
23311
Scratch https://docs.rc.fas.harvard.edu/kb/policy-scratch/ Wed, 13 Dec 2017 10:36:16 +0000 https://www.rc.fas.harvard.edu/?page_id=17408 RC maintains a large, shared temporary scratch filesystem for general use for high input/output jobs at /n/netscratch.

Scratch Policy

Each lab is allotted 50TB of scratch space for its use in their jobs. This is temporary high-performance space and files older than 90 days will be deleted through a periodic purge process. This purge can run at any time, especially if scratch is getting full and is also often run at the start of the month during our monthly maintenance period.

There is no charge to labs for netscratch, but please note that it intended as volatile, temporary scratch space for transient data and is not backed up. If your lab has concerns or needs regarding scratch space or usage, please contact FASRC to discuss.

Modifying file times (via touch or other process) when initially placing data in scratch is allowed, however doing so subsequently to avoid deletion is an abuse of the filesystem and will result in administrative action from FASRC. To reiterate, you may initially modify the file date(s) on new data so that it is not in the past, but should not modify it further.  If you have longer-term needs, please contact us to discuss options.

Tip: Extracting tar archives on netscratch with default options (e.g. tar -xvf) will use the archived timestamps on files, which are often in the distant past and will cause the files to be purged prematurely. To update modified time to the current time on extracted files, add the -m or --touch option to your tar command.


Networked, shared netscratch

The cluster has storage built specifically for high-performance temporary use. You can create your own folder inside the folder of your lab group. If that doesn’t exist or you do not have write access, contact us.

IMPORANT: netscratch is temporary scratch space and has a strict retention policy. 

Size limit 4 Pb total, 50TB max. per group, 100M inodes
Availability All cluster nodes.
Cannot be mounted on desktops/laptops.
Backup NOT backed up
Retention policy 90 day retention policy. Deletions are run during the cluster maintenance window.
Performance High: Appropriate for I/O intensive jobs

 

 

 

 

 

 

/n/netscratch is short-term, volatile, shared scratch space for large data analysis projects.

The /n/netscratch filesystem is managed by the VAST parallel file system and provides excellent performance for HPC environments. This file system can be used for data intensive computation, but must be considered a temporary store. Files are not backed up and will be removed after 90 days. There is a 50TB total usage limit per group.

Large data analysis jobs that would fill your 100 Gb of home space can be run from this volume. Once analysis has been completed, however, data you wish to retain must be moved elsewhere (lab storage, etc.). The retention policy will remove data from scratch storage after 90 days.


Local (per node), shared scratch storage

Each node contains a disk partition, /scratch, also known as the local scratch that is useful for storing large temporary files created while an application is running.

IMPORTANT: Local scratch is highly volatile and should not be expected to persist beyond job duration.

Size limit Variable (200-300GB total typical). See actual limits per partition.
Availability Node only.
Cannot be mounted on desktops/laptops.
Backup Not backed up
Retention policy Not retained – Highly Volatile
Performance High: Suited for limited I/O intensive jobs

 

 

 

 

 

 

The /scratch volumes are a directly connected (and therefore, fast) to temporary storage location that is local to the compute node. Many high performance computing applications use temporary files that go to /tmp by default. On the cluster we have pointed /tmp to /scratch. Network-attached storage, like home directories, is slow compared to disks directly connected to the compute node. If you can direct your application to use /scratch for temporary files, you can gain significant performance improvements and ensure that large files can be supported.

Though there are /scratch directories available to each compute node, they are not the same volume. The storage is specific to the host and is not shared. For details on the /scratch size available on the host belonging to a given partition, see the last column of the table on Slurm Partitions. Files written to /scratch from holy2a18206, for example, are only visible on that host. /scratch should only be used for temporary files written and removed during the running of a process. Although a ‘scratch cleaner’ does run hourly, we ask that at the end of your job you delete the files that you’ve created.

$SCRATCH VARIABLE

A global variable called $SCRATCH exists on the FASRC Cannon and FASSE clusters which allows scripts and jobs to point to a specific directory in scratch regardless of any changes to the name or path of the top-level scratch filesystem. This variable currently points to /n/netscratch so, for example, one could use the path $SCRATCH/jharvard_lab/Lab/jsmith in a job script. This will have the added benefit of allowing us to change scratch systems at any time without your having to modify your jobs/scripts.

]]>
17408