Our lab uses Teams, the Ohio Supercomputer Center (OSC), GitHub, and the T drive to organize our work. Each location serves a different purpose. Keeping files in the right place makes it easier to find what we need, reproduce analyses, and work together as a project moves toward a manuscript. This page expands the “locations of data” item on the onboarding checklist.

One rule applies throughout: protected health information (PHI) must never be stored in GitHub. PHI-containing human data are processed on PDE0023, and a restricted-access folder on the T drive will hold PHI-containing clinical data. Only PHI-free outputs belong in the project’s GitHub repository or on PAS1695.

While the PHI rules are strict, everything else should be considered a convention. We’ve found it works well for our typical manuscript structure, but we’re open to changes. If this is your first project, please try to follow the convention before suggesting alternatives; most of choices reflected here are viable solutions to a problem we experienced.

Teams: Project Materials and Shared Lab Resources

Teams is the home for shared project materials, conference materials, and lab methods.

  • projects: Organize folders by project code. For this purpose, we consider a project to be a manuscript. Every project in Teams should have a corresponding GitHub repository.
  • conferences: Keep posters, abstracts, and talks here.
  • SOPs: This is the master location for lab methods. A typical workflow starts with an SOP from this folder: make a copy, move it into the project directory, and customize that copy for the project. If you notice that the master SOP is out of date, please update it so everyone benefits from the correction.
  • video-tutorials: Keep recordings of challenging workflows that benefit from being demonstrated on video here.

OSC: Software, Large Files, and Human Data Processing

Use the appropriate OSC location for the type of work you are doing:

  • Home directory: Keep GitHub repository checkouts and custom software installations here only. Check the lab’s conda environments for common tools before installing your own copy.
  • PAS1695: Store large, PHI-free files here. Scripts maintained in GitHub that need these large files should read them from PAS1695.
  • PDE0023: Use this allocation to process PHI-containing human data. Store the PHI-free processed outputs in the project’s GitHub repository when they are small enough, or on PAS1695 when they are very large. If you find scripts with hard-coded paths to PDE0023, please let Dan know so we can change them.

GitHub: Scripts and Reproducible Analyses

Every project in Teams should have a corresponding GitHub repository. Keep all scripts in GitHub, along with PHI-free input data files smaller than 100 MB. Never put PHI in a repository. If you are unsure where something belongs, ask in the project Slack channel; how we use Slack is the usual place for those questions.

Use paths relative to the repository for files stored within it. Large external inputs can be read from PAS1695, and clinical source data can be read from the T drive through the processing workflow described below.

Most work begins in an exploratory directory with three subfolders: data, scripts, and figures. Once we identify which analyses will go into the manuscript, create a manuscript directory with the same structure:

project-repo/
├── exploratory/
│   ├── data/
│   ├── scripts/
│   └── figures/
└── manuscript/
    ├── data/
    ├── scripts/
    └── figures/

Copy the relevant files from exploratory into manuscript. This preserves the exploratory work while giving the manuscript analyses a clear home, without requiring us to clean out the exploratory directory or delete analyses that did not make it into the paper. The size and PHI rules apply to both directories.

In cases where the repository is also used for grants, conferences, or other activities, we create a dissemination folder, in which manuscript becomes a subfolder, as are conferences and grants.

project-repo/
├── exploratory/
│   ├── data/
│   ├── scripts/
│   └── figures/
└── dissemination/
    └── manuscript/
        ├── data/
        ├── scripts/
        └── figures/
    └── grants/
        └── NCI-R01_26/
            ├── data/
            ├── scripts/
            └── figures/
    └── conferences/
        └── AACR26/
            ├── data/
            ├── scripts/
            └── figures/
    

While this leads to duplication of files, it ensures that data provenance is maintained for earlier analyses even if source data are updated.

T Drive: Clinical Source Data and Large Data Archives

A restricted-access folder on the T drive will hold PHI-containing clinical data. We also use the T drive to archive large data files.

Scripts can read from the T drive, but keep these dependencies to a minimum. For clinical data, the typical workflow is:

  1. A script called 00_clinical-processing.Rmd, maintained in GitHub, reads the clinical source data from the restricted-access T drive folder.
  2. That script produces a cleaned, PHI-free clinical output, which is stored in the repository when small enough.
  3. All other scripts that use clinical data read the cleaned output from the repository.

This gives downstream analyses a shared clinical dataset and keeps access to PHI-containing source files concentrated in the initial processing step. If the PHI-free output is too large for the repository, store it on PAS1695 and have downstream scripts read it there.