02 / The process
What we do, and what becomes readable
This is not an analysis service. It is instruction in the technical procedure — and the thinking behind it — for reading your own data yourself. Your data stays with you.
Assumptions
This is written mainly for people who have already had whole-genome sequencing done and hold the data. If you have not, I can start further back — with which service to use and what to ask them for.
You will need storage and a connection that can handle files in the tens of gigabytes, along with a working environment on your own machine. Setting that up is part of what we cover.
Stages
-
01
Get the raw data out
How to obtain your FASTQ or VCF from whichever service did the sequencing — GeneLife, Sequencing.com and the rest. Which service releases what format, and how far they will go, is not consistent, so we settle this first.
-
02
Establish what the data is
Which reference genome it was called against (GRCh37 or GRCh38), the coverage depth, the distribution of what was never read. Before any reading begins we fix what this particular data can and cannot support. That includes the regions short-read sequencing cannot reach in principle.
-
03
Pin each variant down
One position carries several names at once: an rsID, HGVS notation, genomic coordinates. On top of that, mixing up the plus and minus strand will flip a genotype without warning. You learn to normalise the notation and to catch the mix-ups.
-
04
Annotate
Which reference databases to consult, and in what order. Allele frequencies shift substantially between populations, so this includes the habit of checking Japanese population frequencies alongside the rest. Judging the strength of the evidence belongs here too.
-
05
Place it on the metabolic map
The genotype, and the phenotype mechanically determined from it, arranged as the structure of a metabolic pathway. The conversion table used is always cited alongside.
Also covered
- Storing the data
- Where to keep your data and how. The minimum settings if you use cloud storage, and the places where accidents tend to happen. Including when not to use the cloud at all.
- Where AI helps, and how much to hand it
- Breaking down terminology, assembling commands, getting the gist of a paper. Generative AI genuinely helps with this work, so we start from how to use it well. Alongside that: what to do about invented rsIDs and invented papers, and how to confirm results against primary databases. And the material you need in order to decide for yourself how much to hand to an outside service. Passing over the data file itself and asking about a particular rsID are two different things.
What I do not take on
- Diagnosis, or any indication of the likelihood of developing a disease
- Dietary, exercise or supplement advice based on your individual results
- Running the analysis on your behalf, or taking custody of your data
- Questioning you about symptoms, condition or medication
Conversation naturally arises while the work is underway. What is delivered, however, is the genotype, the phenotype mechanically determined from it, and the diagram those are placed on. No assessment of your state of health is made.
Fees
Which stages are needed depends on the format of your data and on how far you want to go, so fees are quoted individually.
How to start
To begin, tell me what format your data is in (FASTQ, VCF or otherwise), which testing service produced it, and how far you want to take the reading. I will come back with the stages involved and an approximate quote. You can decide whether to proceed from there.
You will not be asked to send the genomic data itself — not now, and not later. The format and the provenance are all I need.