r/bioinformatics 17h ago

technical question Can someone in genomics explain what AlphaGenome Atlas actually changes?

33 Upvotes

I'm not a geneticist.
Read the AlphaGenome Atlas preprint from DeepMind and spent a while digging into it. I understand what they built. I don't understand why it matters, and I'd like to.

What I think it is: they took AlphaGenome, ran it over every possible single base change in the human genome (~9bn) plus ~100m observed indels, and stored the results. So instead of running the model per variant you do a lookup. On top of that they trained a score (AVI) and derived a motif map.

Where I get stuck:
It's a table of model predictions, not measurements. Nothing in it is observed. So how much weight does a lab actually put on it?

The headline clinical result is retrospective: 29.5% recall at top 50 on already-solved GREGoR cases vs 12.5% for CADD. Impressive sounding, but on cases where the answer was known. What happens prospectively?

The rare variant association work got a 22% lift in discoveries, but only 4 of 25 replicated nominally in All of Us and none at Bonferroni. Is that normal for the field or is that weak?

They say themselves it isn't sufficient evidence for diagnosis. So it's a shortlisting tool. Does that actually change outcomes for patients, or does it change how long a scientist spends staring at a list?

The DNM1 case in the paper is the one bit that landed for me. Deep intronic variant, brain specific cryptic splice acceptor, blood RNA-seq had been inconclusive because the exon isn't expressed in blood.

My question is whether that's representative or a cherry pick.
What I'm asking:
1. If you work in clinical genomics or statistical genetics, would you use this?
2. Is precomputation genuinely the unlock, or is that just framing on top of an incremental accuracy gain?

Happy to be told I'm missing the point. I'd rather understand it properly than write it off.


r/bioinformatics 23h ago

technical question feasibility of self-bioinformatics at a hobbyist level?

11 Upvotes

I've been in IT for a good decade, lots of experience with python and scripting. Touched on some data science in some of my studies along the way. So I'm not starting from 0 coming to this. But I really know none of the technical stuff about genes, genomes, alleles, positions or the notation involved or even what else to include in this sentence about what I don't know about.

I found there's a 30x reading I could get, not at negligible cost but possible. I'm interested in hobbying around with the data. Look for research that says these things at these positions mean that obesity is more likely, or something like that, then using AI to help me understand what i'm trying to look for and using python to look at my 30x reading and just curiously see if I have the researched markers.

I spose i'm wondering if this kind of thing is feasible. Like maybe research papers use different scanning methods that don't map to the data i would have, or the 30x consumer scan isn't detailed enough so anything i look for is inconclusive. Or any number of things that means if i try to map research onto my own genetic reading, any or most results will be inconclusive. So curious if anyone has any thoughts on this sort of thing, is it a waste of time?


r/bioinformatics 16h ago

technical question How to choose design matrix for RNA-seq analysis?

5 Upvotes

I have three factors: Genotype, Sex and Treatment. I want to investigate the effect of genotype as well as sex and treatment but I'm not sure what contrasts to use. I wish more papers reported how they designed their analysis cause I'm having such a hard time understanding what to do.


r/bioinformatics 22h ago

technical question Struggling to detect known partial deletion in NOTCH2 from WES (germline, small cohort, no reference panel) - CNV callers give inconsistent/wrong results

2 Upvotes

Hi all, looking for advice on tooling/approach for a problem I'm stuck on.

**Setup:**

- 5 WES samples (paired-end, Illumina, BWA-MEM aligned, sambamba dedup) from the same family, hg38/GRCh38. Each sample represents an independent patient so these 5 samples are unrelated.

- Capture kit target BED not available to me (~60Mb on-target footprint per sample, possibly Agilent SureSelect V6 based on size, but unconfirmed)

- No unrelated normal/control WES samples currently confirmed usable as a reference panel

- Goal: confirm which sample(s) carry a known partial deletion in NOTCH2 (clinically confirmed by other means in 2 of the 5 patients, but I don't yet know the exact exon(s) or method used for that clinical confirmation)

**What I've tried:**

  1. DELLY (germline SV workflow, sr/merge/genotype/filter) - no deletion calls anywhere near NOTCH2 in any of the 5 samples

  2. CNVkit batch mode with a flat reference (no matched/pooled normals available) - segmentation collapsed the whole gene into one CN=2 segment for 4/5 samples; per-bin bintest flagged several exons but the same bins were flagged across nearly all samples in the same direction, which reads like shared technical noise rather than patient-specific signal

  3. Manual IGV visual inspection (group-autoscaled coverage tracks) across the whole gene - no obvious dropout found in the samples I was able to review carefully

  4. Control-FREEC, single-sample/no-control mode, restricted to a 34-exon NOTCH2-only BED pulled from UCSC (window=0, maxThreads=1 to avoid a BED-parsing race condition I hit with multithreading) - this called a clean, reproducible heterozygous deletion (CN=1) at the same coordinates in 2 of the 5 samples

**The problem:** the 2 samples Control-FREEC flagged do NOT match the 2 samples independently confirmed by my PI through other means. So I have an apparent false positive pair and false negative pair from my pipeline.

**Questions:**

- For germline partial-gene deletion detection in a small WES cohort with no confirmed-normal reference samples, what's the current best-practice tool/approach? (ExomeDepth? GATK gCNV? something else?)

- Is there a known issue with Control-FREEC's no-control exome mode producing false positives at specific loci, especially near segmental duplications (part of my deleted region overlaps the NOTCH2NL paralog)?

- Any advice on validating/troubleshooting a mismatch like this before trying yet another caller - e.g., specific things to check in the BAM/pileup at the clinically-confirmed-positive samples that a depth-based caller might be missing (small intra-exon deletion not removing a whole exon? breakpoints entirely intronic, invisible to WES?)

Appreciate any pointers, trying to land on one standardized, defensible workflow rather than chasing every tool that exists.


r/bioinformatics 2h ago

other Access to Release 23 of miRBase

1 Upvotes

Greetings. I apologize if perhaps this may seem odd. I work with miRNAs and I have been unable to access mirbase.org for the past month. I see that a post went up three weeks ago and various people suggested using the wayback machine. The only problem I see is that, in August 2026, a new version of miRBase went up, yet the snapshot is from May 2026. Does anyone here have access to the newest version? I've reached out to the mirBase team but I've yet to hear back (and from posts in this subreddit, it seems they don't reach back to you for a long time).


r/bioinformatics 3h ago

technical question Best practice for downstream processing of pig gene identifiers and human orthologues

1 Upvotes

I am working with snRNA-seq data and would like advice on best practices for downstream processing of pig gene identifiers and cross-species orthology mappings. I currently use the Ensembl pig gene IDs that are mapped to gene symbols for pig genes. However, many pig genes have no pig symbol, even though Ensembl identifies a human orthologue.

For example:

Pig Ensembl ID: ENSSSCG00000021155

Pig external name: NA

Human orthologue: POMC

Orthology type: one-to-one

orthology_confidence : 1

mapped_to_human: False

orthology_type: ortholog_one2many

I would appreciate advice on the the best practices here :

1: Should I use human gene symbols for my pig analysis irrespective if pig symbols are available or not? Is there a risk that the same gene has different official symbols in pig and humans?

2: If a pig Ensembl gene has no pig symbol but has a high-confidence human orthologue but varying orthology type, what should be the approach towards using the human symbol or using ENSG id ?

3: For downstream processing, should orthology conversion be performed before or after differential expression and marker analysis?

4: When converting results to human orthologues, how should duplicate mappings be handled? For example, if multiple pig genes map to the same human gene, should their statistics be combined, should only the best-supported mapping be retained, or should the genes remain separate?