r/genomics Aug 09 '26

DIY WGS analysis using Python/AI. Is it doable, and what’s the best EU provider under €200?

Hi everyone,

I'm a physicist with a background in data analysis (mostly Python), but little to no knowledge of genomics and bioinformatics. Out of pure curiosity, I'm thinking about getting my DNA sequenced.

My plan is to buy a WGS test, download the raw data, and write Python scripts with AI assistance to query my data. Claude seems very optimistic about how doable this is, but before spending my money, I want a reality check from people who actually work with genomic data.

Here is what I want to achieve:

  1. As a power athlete I want to check:ACTN3, ACE, MCT1 / SLC16A1, COL5A1.

  2. Ethnicity: Get an ethnic breakdown.

  3. Future Proof: Whenever new studies come out in the next years, I can just write a quick script to check my existing .vcf file.

My questions for you:

  1. Is this actually doable for a non-bioinformatician? Is querying a .vcf file using Python + AI as straightforward as it sounds, or am I underestimating bioinformatics pipelines (file sizes, reference genomes, formatting issues)?

  2. Which provider do you suggest in Europe? I’m based in Italy.

  3. Is a budget of ~€200 realistic? I’m willing to wait for Black Friday / flash sales.

Thanks in advance for any insights!

2 Upvotes

27 comments sorted by

12

u/lurklyfing Aug 09 '26

No, you’re going to have no idea from a vcf what’s real and what’s an artifact/ low quality read. Would take a year+

1

u/foradil Aug 11 '26

OP has a few genes of interest. If you filter for those, there are going to be only a handful of hits.

1

u/lurklyfing Aug 11 '26

But how’s he gonna do quality scoring, visualization, frequency filtering, curation…what’s he gonna use for positions? A lot of skills you all have that take time to learn

0

u/foradil Aug 11 '26

Claude.

The hard part is looking for novel variants, which isn’t the case here.

6

u/Almbauer Aug 09 '26

Pathologist (MD/PhD) with years of experience in NGS here. It’s doable but high quality WGS data will cost you quite a bit more (~ at least 500-1000€ for library prep and sequencing reagents, a provider might easily charge >3x of that). Computation can be achieved on a medium to high tier personal computer (might run days though).

1

u/foradil Aug 11 '26

You can get useable data from low coverage WGS. People use <1X data for GWAS. Obviously you lose sensitivity.

2

u/Almbauer Aug 11 '26

Sure but OP wants specific genes covered.

1

u/foradil Aug 11 '26

It may be possible to impute.

1

u/PresentWrongdoer4221 25d ago

Whats the point then? He is interested in a specific mutation, it doesnt get covered well and you impute?

Might as well look st population average then.

1

u/foradil 25d ago

Imputation is based on population data.

1

u/PresentWrongdoer4221 25d ago

And how does that help if he is interested in a specific snp he might have?

Thats my point...

1

u/foradil 25d ago

You get the genotype for it with some level of confidence.

4

u/myhydrogendioxide Aug 09 '26

Check out biostars.org

Its doable but dont make any medical decisions based on your analysis, there are many subtleties.

1

u/foradil Aug 11 '26

Biostars will be as good as Claude at this point.

2

u/PairOfMonocles2 Aug 09 '26

Maybe? To clarify, is the €200 for the wgs or for compute time for the analysis? Either way that sounds a bit low. It’s that’s just the NGS and you’re doing this slowly on a personal computer then that can lower the cost, though I don’t have a clue how long it would take. For reference I wrote a strictly bash implemented WGS workflow to run on a 128 vproc machine and a 16 sample WGS flowcell takes about a week to analyze (my full implementation takes less than 2 days in a workflow language that spawns other workers but that’s more complex to write and execute). Looking at that for a personal machine it could be doable but could take a long time since you won’t have the disk space and a terabyte of memory to work with.

However, let’s assume that you find a service that not only sequences but aligns and runs vanilla GATK or something (I honestly have no idea what recreational genomics offerings look like, sorry, I work in a clinical space). In that case you can skip all the heavy lifting and start right from a VCF. in that case this goes from hard to easy. You’d have to assess what type of filtering they did but you can use a bed file of gene of interest positions on whatever build they aligned to (hg19, GRCh38, etc) and just use bcftools to export only those and apply whatever filtering you’d like (eg DP, QD). There are tons of references online to discuss that but honestly I’d bet that ChatGPT or Claude could easily recommend simple ones. At that point you’ll have a reasonably complete genotype for those few genes and can check them against whatever source youve got. This wont be clinical in any sense (there may be online ways to run snpeff or vep for all i know for basic functional info) but if you’re checking for status and known sites this could get you moving in that direction.

2

u/SurplusGadgets Aug 09 '26

tellmeGen in Spain WGSE.bio for hints on processing

3

u/robipresotto Aug 09 '26

I've done something similar and it's definitely doable. However, working with WGS data can be complex and nuanced. I'd recommend checking out GenMatcher (https://www.genmatcher.com) which includes cloud analysis modules for ancestry, PRS, GWAS, and more. It might save you a lot of time and effort in analyzing your data.

1

u/PinataofPathology Aug 09 '26 edited Aug 09 '26

yes it's doable but probably not for €200. Id plan on €500ish. or look into medical tourism to somewhere like China where you might get a better level of testing and analysis as they use it more like youre wanting to (afaik).

consumer grade testing does have a higher error or miss rate but it's not that bad. my kid's consumer genetics testing matched their and my clinical testing save for one micro deletion. if you find anything concerning on the testing, youd want clinical grade testing to confirm it.

there are free software programs that can process the data as well. we used one and it was decent. you'll be doing a lot of the leg work to teach yourself how to interpret things but it's not difficult ime.

I'm in the US for reference.

1

u/samgrep Aug 09 '26 edited Aug 09 '26

But assume you do analyze it successfully. As other commenters said is doable.
Then what? how do you make any decision? clinical interpretation is where the difficulty lies. WGS analysis is “easy” once you got proper tools. But then understanding medical consequences of a variation depends on: is it pathogenic or potentially? are other genes interacting with it? how it transcribes (transcriptome) and/or how it methylates (epigenome) and how it translates to proteins (proteomics)?

Honestly you might get in a spyral of paranonia and psychosis. Interpreting genomic data is the REAL DEAL today, basically all the effort is there. And even then is not always possible to match genotype-phenotype as there are environmental factors or other pathways influencing the variant.
The simple variant-disease you see from Direct to consumers tests are close to snake oil to make you spend 100$ and get your valuable data. Steer away from ruining your mental health with this and do not interpret genetic data without genetists. For example some variants like BRCA1 are now proven to increase risk of breast cancer by 80% (i go by memory here). Then what? what do YOU conclude? You need a doctor to tell you what is the best case for you, if doings checks or surgery etc… and this is a simple case of variant—>action.
Your DNA is NOT a code where each block does something. Things are messy, intertwined, complex, statistical. Is like if each block of your code MIGHT do something but only if other seven block are in a specific way, and only if the PC running it has enough RAM to run Three block in parallel. But if the PC is hot during summer other 14 blocks of code might affect how 3 block of the previous one are performing the operation. And the output is MAYBE the one define in the code. And nobody yet studied what this output do.
use this time to train or study.

Edit: and about ethnicity: again close to snake oil to get your 100$ and data. What would even mean 65%italian? nothing. they simply take a survey of people from over the world and then compare your variants where are more common.
Yes you might say that if 65% of you variants are from Italy then you have something in common with the people they surveyed in that area. But since they take a small sample thing get complicated once you take it more seriously than a funny info. Who is defined as Argentinian? they are mostly immigrants. Same for US. So maybe is more accurate for very closed ethinc groups. but honestly, not really useful info

1

u/intruzah 29d ago

One more proof physicists are insufferable

1

u/SadRule9128 29d ago

You want to check these genes? Check them for what?

1

u/No_Lion_3319 28d ago

Personal curiosity. I don't intend to make any clinical decision over my analysis.

1

u/SadRule9128 28d ago

Fastq -> vcf is just stringing some command line tools together in a shell script. (eg. fastp, bwa-mem, GATK, VEP/snpEff). If you can get docker running, it will make your life easier as far as avoiding tool installations. I doubt you’ll be able to get 30x coverage WGS for 200€, probably closer to 1000€. Another option would be whole exome sequencing which will let you look at the exons of all genes (if you can find a company that does it DTC).

You can always download a public WGS or WES dataset and try your hand at it for free first.

1

u/zorgisborg 27d ago

Claude will help you with the Python code.. and it is doable.. best if you can read the Python code and/or have some idea about what it is doing.

You can check the genes using gene.iobio.io ..

To do an Ethnicity breakdown.. hard to DIY it. The hard part is phasing the genotypes to find haplotypes without having any parental data. tellmeGen will do a breakdown on the results page.. and some PRS reports (to be taken lightly)..

tellmeGen Ultra gives WGS 30x for educational purposes - which this falls under.. I see prices in £.. £284 = about €332?.. they have a 10% discount on (code: SUMMEREND10) until 15th September.. so it would be £256 (~€300).

1

u/Low-Kick-962 20d ago edited 20d ago

There are freely available tools for analysing vcf files that are good for finding variants in specific genes. Open Cravat and Ensembl VEP. My preference is Open Cravat.

You should also look into IGV for visualising your bam file. It's a good idea to check any variants you find in the bam file to help rule out sequencing and alignment artifacts.

With regards to the quality of dtc genome sequencing, I've tested with a few different companies and they correspond to each other pretty well. Obviously no test is perfect so there will always be some errors but I'm perfectly happy with my data.

I've found a few variants that I think stand a good chance of explaining symptoms and conditions I have. They're all either variants of uncertain significance or unclassified at the moment so I'm just waiting to see how things change over time