“We have results from our clinical trial, and need Diamond Age’s help in analyzing the data.” These words are like music to my ears. Clinical trial projects are exciting to me because they are some of the first looks we have at how treatments behave in human patients. We get to see if our hypotheses from the lab hold up in the clinic and start to assess which patients the drug works for, and how. From the outside, the core activities – omics analysis, NGS data processing, biomarker discovery – might appear to be the same as their pre-clinical counterparts. However, there are important differences these types of analysis bring, both in approach and mindset, that we need to be aware of before embarking on a project.
The Allure and Limits of the Controlled Model
Coming from a computational biology background working on data solely from humans, my first mouse experiment was mind blowing. You mean we have biological replicates? All samples have the same genetics and same disease state?! The experiment design is fully factored with appropriate controls?!?! These are luxuries that we rarely see working with human data. However, the more experience I gained working with models, the more I understood the uncertainty of how what we found would translate to human patients. Every result was a nudge of information about how a disease or therapy works in humans. I think it’s important we remember that each of these experiments is not an end in itself, but a means to discover something about human biology and disease. Models are, after all, imperfect proxies for what they attempt to be modeling.
The Reality of Human Variation
The moment we transition to clinical data, we are no longer dealing with a simple model; we are engaging with the full complexity of the human experience. Every individual patient carries their own patient-specific confounding factors such as genetics, lifestyle, concurrent medications, or other diseases, which muddles the biological signal we are looking for. Furthermore, even within a single patient, visit-to-visit variability—like having a common cold or a minor inflammatory event—adds to the data’s complexity. Instead of relying on model organisms, we are sampling patients from the real world. This increased variability is a key feature of the dataset, even when it makes interpretation of the results difficult.
The Metadata Challenge
Let’s be honest – metadata is always a challenge. Even with tidier preclinical studies, sample annotation can be messy, inconsistent, and not machine readable. This challenge is multiplied when data collection and annotation are outsourced to multiple clinical sites and CROs, and having a greater number of variables to track (dose, sex, age, disease subtype, etc). It’s imperative to take extra care with data wrangling and cleaning.
Once our sample annotation is cleaned and harmonized, our work here isn’t over. We must take the time to diligently look for and correct confounders in the data, even with covariates we don’t expect to have an impact on the data. For example, I’ve seen an instance where one clinical site collected or stored the samples differently, and those samples were clear outlier clusters. This step is both more critical and inherently harder to achieve with clinical data.
Redefining ‘Signal’ in Low-Power Scenarios
For bioinformaticians, the most crucial pivot is changing how we define a valuable “signal.” In well-designed pre-clinical studies, best practices analyses can lead to statistically significant results. In small Phase 1 and 2 studies, chasing statistical significance can be a fool’s errand due to low statistical power. Small n, multiple dose cohorts, or different indications can force data stratification before analysis. I’ve seen Phase I “bucket” trials that have a dozen or so patients, include a range of cancer types, and have a handful of responders (also distributed among different cancer types), and have been tasked with finding transcriptional biomarkers of response in this dataset. With such a small sample size, we get to flex our creative muscles on how to look for the relevant signal.
At Diamond Age we have found success in looking beyond low p-values to identifying plausible biological patterns. Sometimes the focus of the analysis shifts to counting and pattern-finding instead of inferential statistics. A modest but consistent biomarker shift across a small, defined subset of patients, even if statistically non-significant against the whole cohort, can be a high-value lead. It requires biological astuteness to identify which of these patterns could be relevant – even if unexpected. We also need to be able to communicate these results to the clinical team, and work hand in hand with them to refine methods and interpretations. Finding these biologically relevant signals often takes perseverance and creativity, but I have seen these small insights turn into high-impact scientific paths, even leading to novel target programs.
Honoring the Weight of the Data
Every single observation in our dataset represents a real patient visit and a real step in their treatment journey. That high variance we’re modeling is not just noise; it’s a reflection of an individual’s unique characteristics. Each sample we analyze is connected to a person who volunteered for a trial, hoping for a better outcome, and trusting the research process.
This data carries an emotional weight. Our expertise is what translates these individual patient contributions into actionable, population-level knowledge. We must apply the most robust and rigorous methods possible to ensure we extract every last ounce of usable insight. When we commit to understanding the messy metadata and diligently digging for patterns in low-N cohorts, we aren’t just doing good science; we are fulfilling trust placed in us. Our dedication to high-quality, pragmatic analysis is how we do justice to the patients who are, quite literally, giving us their data for the benefit of others.
Pragmatic Action for the Analyst
So, what does this fundamental shift mean for you, the analyst?
The good news is that you don’t necessarily need an entirely new set of bioinformatics skills. The core competencies remain essential. What you do need is a heightened awareness of the additional complexities.
- Engage with the clinical team about the nuances of the study design, sample collection variables, and potential confounders. Just like any good analysis, the results are only as good as the context you bring to the data.
- Consider pivoting your methodology. Instead of rigidly filtering for a strict p_adj < 0.05, try sorting by significance, experimenting with different visualizations, and looking for patterns within pre-determined, biologically relevant gene sets.
- Be prepared to dig a few layers deeper, combining computational skill with biological astuteness, to find that interesting, quiet biological signal.
Working with clinical trial data might be a little messy and difficult but know that whatever you discover will have true, translatable value that directly informs future trials and benefits patients. That’s the ultimate win.