Bioinformatics Coding Guide

Learn one step at a time using the lab’s workshop notebooks.

Single-cell and spatial transcriptomics · Analysis workflow and R coding · Merged workflow v3 · FASTQ → Cell Ranger → R

This guide explains analysis code adapted from the workshop notebooks. It does not run R or provide live AI chat. Read the source review before running advanced sections. Choose where to run R yourself; this guide focuses on analysis order, code changes and understanding results.
LSUHS computing resources · Spartacus HPC

LSU Health Shreveport — High-Performance Computing Core

Spartacus HPC is the institution’s high-performance computing cluster for genomics and other data-intensive research. The Core provides browser-based access, with software pre-installed or available by request.

Official Spartacus HPC information →

Contact: Jarrod Sawyer · Project Coordinator, Research Data Team
Jarrod.Sawyer@lsuhs.edu

This is the information page, not a verified login portal. Ask the Core about accounts, available R software and current access arrangements. You can follow this coding guide in another R environment as well.

Use your own data · biological replicates · when to combine samples

What is an example?

The source workshop uses human cancer pre/post samples, named preT and postT. GSM/SRR accessions, ITGB1 and other genes, Cancer cells / Memory CD8+ T cells labels, prostate tissue domains, file paths, QC cutoffs and plot interpretations are example-specific. Use your own sample names, organism, marker genes, validated annotations and actual results. Do not copy the workshop biological conclusions into your experiment.

Record the experimental design before mapping

Use a sample sheet. library_id identifies a sequencing library/GEM well; sample_id identifies the biological specimen; donor_id identifies the person, animal or other relevant biological unit; condition records treatment; batch records preparation/processing batch. The downloadable six-row sheet is a fictional paired example with three donors, not data from the source workshop and not a universal required sample size.

Keep biological replicates separate during mapping and QC

Run Cell Ranger for each appropriate library/GEM well. Multiple sequencing lanes/runs from the same GEM well can be inputs to the same count run; different biological specimens must not be combined into one apparent sample. For multiplexed libraries use the assay-specific demultiplexing workflow and recover per-specimen identities before downstream analysis. Perform ambient correction and QC independently per relevant library/sample, then retain specimen and donor metadata.

Combine objects after per-sample processing

For this SCT workflow, normalize each QC-filtered object with SCTransform, then supply the complete object list to integration before joint PCA, neighbor finding, clustering and annotation. The workshop list(preT, postT) is only a two-object example. Rename cell barcodes with unique library IDs and preserve sample_id, donor_id, condition and batch. merge concatenates objects; integration adjusts a representation and is a separate choice, not mandatory for every study.

Use replicate-level inference for condition comparisons

After validating cell labels, select one cell type and aggregate count input by biological specimen (pseudobulk). Keep specimens separate: do not pool all controls into one column. An unpaired experiment compares independent biological units; a paired experiment tracks the same donor across conditions and includes donor in its design. Replicate sufficiency and power depend on the study. The workshop two-sample example cannot establish replicated group-level effects.

Technical repeats and repeated measures

Multiple runs of the same GEM well do not increase biological n. Multiple GEM wells or technical libraries from the same specimen need distinct cell IDs and remain one specimen-level biological observation for appropriate aggregation. Multiple timepoints from one donor are paired/repeated measures; use a matching design rather than treating them as independent donors. Multiple spatial sections from one individual likewise do not automatically become independent biological replicates.

Do not lose the count input

Use appropriate nonnegative count data for count-based sample-level testing, not integrated values, scaled values or averaged expression. Original analytical RNA counts and corrected fractional values are not interchangeable. The source SoupX correction can produce fractional values; the provided pseudobulk template stops on these so a reviewed count-source choice is required. Inspect cell abundance and missing populations at the sample level.

Fictional paired sample sheet

library_id,sample_id,donor_id,condition,batch,counts_rds
L01,S01,D01,control,B1,processed/S01_sct.rds
L02,S02,D01,treated,B1,processed/S02_sct.rds
L03,S03,D02,control,B1,processed/S03_sct.rds
L04,S04,D02,treated,B1,processed/S04_sct.rds
L05,S05,D03,control,B2,processed/S05_sct.rds
L06,S06,D03,treated,B2,processed/S06_sct.rds

Save a copy as sample-sheet.csv and replace every row with your actual libraries. counts_rds points to independently processed SCT objects; those files are not created by this example CSV.

Own-data integration template · replaces the two-sample integration block
# OWN-DATA TEMPLATE: use instead of list(preT, postT) integration.
# Input: one saved, independently QC'd and SCT-normalized object per library.
# Change these example paths/labels to your actual design.
library(Seurat)
samples <- read.csv("sample-sheet.csv", stringsAsFactors = FALSE)
required <- c("library_id", "sample_id", "donor_id", "condition", "batch", "counts_rds")
stopifnot(all(required %in% names(samples)), nrow(samples) >= 2)
stopifnot(!anyDuplicated(samples$library_id), all(file.exists(samples$counts_rds)))
stopifnot(all(vapply(samples[required], function(x) all(!is.na(x) & nzchar(trimws(x))), logical(1))))
objects <- lapply(seq_len(nrow(samples)), function(i) {
  obj <- readRDS(samples$counts_rds[i])
  stopifnot(inherits(obj, "Seurat"), all(c("RNA", "SCT") %in% Assays(obj)))
  obj <- RenameCells(obj, add.cell.id = samples$library_id[i])
  obj$library_id <- samples$library_id[i]
  obj$sample_id <- samples$sample_id[i]
  obj$donor_id <- samples$donor_id[i]
  obj$condition <- samples$condition[i]
  obj$batch <- samples$batch[i]
  obj
})
names(objects) <- samples$library_id
features <- SelectIntegrationFeatures(object.list = objects, nfeatures = 3000)
prepared <- PrepSCTIntegration(object.list = objects, anchor.features = features)
anchors <- FindIntegrationAnchors(object.list = prepared,
                                 normalization.method = "SCT",
                                 anchor.features = features)
combined <- IntegrateData(anchorset = anchors, normalization.method = "SCT")
# This replaces the two-object integration block; continue PCA/clustering on combined.
# Preserve RNA counts and all metadata. Evaluate batch correction biologically.
print(table(combined$sample_id, combined$condition))
saveRDS(combined, "integrated_with_sample_metadata.rds")
Own-data pseudobulk template · after annotation
# OWN-DATA TEMPLATE: run after cell annotation, NOT on the workshop's
# one-sample-per-condition demo. Requires independent biological replication.
# Add installation separately if absent: BiocManager::install("DESeq2")
library(Seurat)
library(DESeq2)
TARGET_CELLTYPE <- "REPLACE_WITH_YOUR_CELL_LABEL"
PAIRED <- TRUE  # TRUE only when each donor contributes both conditions
stopifnot(TARGET_CELLTYPE != "REPLACE_WITH_YOUR_CELL_LABEL")
# Assign your validated annotations to obj40$celltype before this template.
needed <- c("sample_id", "donor_id", "condition", "celltype")
stopifnot(all(needed %in% colnames(obj40[[]])))
selected <- colnames(obj40)[obj40$celltype == TARGET_CELLTYPE & !is.na(obj40$celltype)]
stopifnot(length(selected) > 0)
ct <- subset(obj40, cells = selected)
# With Seurat v5 split RNA layers, join them before accessing counts.
if (inherits(ct[["RNA"]], "Assay5")) ct <- JoinLayers(ct, assay = "RNA")
raw_counts <- GetAssayData(ct, assay = "RNA", layer = "counts")
# Count-based DE needs valid count input. Fractional ambient-corrected counts
# require an explicit reviewed choice of count source/correction method.
values <- if (inherits(raw_counts, "sparseMatrix")) raw_counts@x else as.vector(raw_counts)
stopifnot(all(is.finite(values)), all(values >= 0))
if (any(abs(values - round(values)) > 1e-8)) {
  stop("RNA counts are fractional: review count source; do not round silently.")
}
meta <- unique(ct[[]][, c("sample_id", "donor_id", "condition"), drop = FALSE])
stopifnot(!anyDuplicated(meta$sample_id))
# Cells are aggregated within each specimen, not across all controls/treatments.
pb <- AggregateExpression(ct, assays = "RNA", group.by = "sample_id",
                          return.seurat = TRUE)
counts <- GetAssayData(pb, assay = "RNA", layer = "counts")
stopifnot(!is.null(pb$sample_id), ncol(counts) == length(pb$sample_id))
idx <- match(pb$sample_id, meta$sample_id)
stopifnot(!anyNA(idx))
coldata <- meta[idx, , drop = FALSE]
rownames(coldata) <- colnames(counts)
coldata$condition <- factor(coldata$condition, levels = c("control", "treated"))
stopifnot(!anyNA(coldata$condition))  # replace factor levels for your real groups
coldata$donor_id <- factor(coldata$donor_id)
if (any(tapply(as.character(coldata$condition), coldata$donor_id,
               function(x) anyDuplicated(x) > 0))) {
  stop("Multiple specimens per donor-condition: review repeated-measures design.")
}
if (PAIRED) {
  donor_table <- table(coldata$donor_id, coldata$condition)
  if (any(donor_table != 1)) stop("Paired template requires one specimen per donor per condition.")
  if (nrow(donor_table) < 2) stop("No replicated donor pairs available.")
  design_formula <- ~ donor_id + condition
} else {
  if (any(table(coldata$donor_id) != 1)) stop("Unpaired template requires independent donors.")
  if (any(table(coldata$condition) < 2)) stop("No biological replication in one or more groups.")
  design_formula <- ~ condition
}
# These checks prevent obvious invalid designs; they are not a power calculation.
# If batch and condition are confounded, a model cannot separate their effects.
mm <- model.matrix(design_formula, data = coldata)
stopifnot(qr(mm)$rank == ncol(mm), nrow(mm) > ncol(mm))
keep <- rowSums(counts >= 10) >= 2  # example low-count filter: review for your design
counts <- counts[keep, , drop = FALSE]
stopifnot(nrow(counts) > 0)
dds <- DESeqDataSetFromMatrix(countData = as.matrix(counts),
                              colData = coldata, design = design_formula)
dds <- DESeq(dds)
result <- results(dds, contrast = c("condition", "treated", "control"))
write.csv(as.data.frame(result), "pseudobulk_treated_vs_control.csv")
# Review sample-level QC, cell numbers per specimen, effect sizes and adjusted p-values.
# Missing cell types are not zero-expression pseudobulk replicates.

Templates have not been run on experimental data. They illustrate the workflow and include design checks; package/version compatibility and study-specific choices still require validation.

Based on the user-supplied Bioinformatics Workshop 2026 notebooks. Workshop source authorship is retained. Explanations and review notes were prepared for Wang Lab.