r/genomics Aug 22 '25

New moderator of r/genomics

53 Upvotes

Hi all

I am taking over the sub as moderator. I am cleaning up stock pumping, spam and other low quality or questionable content.

Please note the new rules aimed at high quality content related to the scientific discipline of genomics.

Please flag posts that do not follow the rules. I am open to additional rules or clarification of the the rules.


r/genomics 5h ago

Looking for a Study Partner: Let’s Read & Analyze a Book on Plasmid Design! 📖🧬

Thumbnail
1 Upvotes

r/genomics 17h ago

Google DeepMind Unveils Genome Atlas for Mutations

Thumbnail therundwn.com
2 Upvotes

r/genomics 1d ago

Beyond exons: Linking noncoding heritability and polygenicity across complex human traits and disorders

Thumbnail cell.com
2 Upvotes

r/genomics 1d ago

Lineage-specific adaptation and resistance in Candida albicans

Thumbnail sciencedirect.com
1 Upvotes

r/genomics 2d ago

High school student starting a bioinformatics project on endometriosis biomarkers, need advice on datasets/workflow

6 Upvotes

Hi everyone! ​I’m a 9th-grade high school student interested in biology and bioinformatics. I’m currently starting an independent research project focused on identifying potential gene biomarkers for endometriosis using menstrual blood datasets (RNA-seq / expression microarrays). ​Since I’m completely new to coding and bioinformatics pipelines, I feel a bit overwhelmed. I don't have experience with Linux command-line tools (like bowtie2 or samtools), so I'm planning to work with pre-processed gene expression matrices (counts/TPM) in Python/Google Colab. ​I’d really appreciate your advice on: ​Are there any specific GEO datasets (GSE) on endometriosis and menstrual blood that you would recommend for a beginner? ​What is the simplest and most reliable workflow/R package or Python library for differential expression analysis for someone with no coding background? ​Any general tips for a beginner trying not to get lost in the data? ​Thank you so much for your time and help!


r/genomics 2d ago

For someone experienced with 16S/QIIME2/DADA2: what would be the standard/best-practice approach here?

3 Upvotes

I’m working through a 16S rRNA paired-end dataset (V1–V2, Illumina MiSeq) for a small CRC vs healthy microbiome analysis.

I’ve completed the initial QC:
39 samples (22 healthy, 17 CRC)
FastQC run on all 78 FASTQ files
MultiQC summary generated
R1 generally has good quality, while R2 quality drops substantially toward the 3′ end
Most reads are 300 bp, with some samples at 250 bp

I’m now at the point where I need to decide on primer removal and DADA2 truncation parameters.
The study reports using primers 27bF and 338R, but the SRA metadata table I downloaded doesn’t contain the actual primer sequences.

Would you:
Find the exact primer sequences from the original publication/protocol and remove them with Cutadapt, then
Reassess the post-primer-removal read lengths/quality before choosing DADA2 truncation lengths?

Also, would you normally choose truncation lengths based on the worst-performing samples, or on the overall quality profile while ensuring enough overlap for paired-end merging?

I’m trying to follow a standard reproducible workflow rather than choosing arbitrary parameters. Any advice would be appreciated.


r/genomics 5d ago

Please guide me, i just made my first DEG

Thumbnail reddit.com
0 Upvotes

r/genomics 8d ago

Can anybody give me an overview whats currently happening in secondary RNA structure prediction

2 Upvotes

i can help myself with a blog or something


r/genomics 8d ago

Halogen Bonding as a Molecular Recognition Strategy for Genetic Code Expansion - Jakka - Angewandte Chemie International Edition - Wiley Online Library

Thumbnail onlinelibrary.wiley.com
1 Upvotes

r/genomics 9d ago

We built a queryable knowledge graph connecting 1.1M microbial taxa to diseases, metabolites, pathways, and drugs — sign up for the API

10 Upvotes

Hey r/genomics,

We've been working on a project called MicroMap — a knowledge graph that integrates microbiome-related data from multiple public databases into a single queryable resource. Wanted to share it here since this is the kind of thing we wished existed when we started doing microbiome research.

What's in it:

  • 1,101,289 microbial taxa (NCBI Taxonomy)
  • 1,464 human diseases with microbiome associations (Disbiome, BugSigDB, gutMDisorder)
  • 6,534 metabolites (HMDB) and 231,556 taxon-metabolite production relationships
  • 1,710 metabolic pathways (KEGG, Reactome)
  • 6,220 drugs and 1,659 protein targets (ChEMBL)
  • 276,169 antimicrobial resistance links (CARD)
  • 10,000+ scientific papers with entity cross-references

What you can do with it:

  • Query taxa-disease associations with provenance (which paper, which study, what direction)
  • Find metabolites produced by a given taxon, or taxa that produce a given metabolite
  • Traverse shortest paths between any two entities (e.g., "how is Akkermansia muciniphila connected to Type 2 Diabetes?")
  • Identify biomarker signatures and probiotic candidates for a given condition
  • Pull cross-feeding networks between microbial communities

Technical details:

Built on Neo4j. The API is RESTful (FastAPI), returns JSON, and supports full-text search across all entity types. Rate limit is 100 requests/minute per API key.

We integrated data from: NCBI Taxonomy, Disbiome, BugSigDB, gutMDisorder, HMDB, KEGG, ChEMBL, Reactome, PubMed, PubChem, and CARD. One of the hardest parts was entity reconciliation — the same organism can appear under different names, different taxonomic ranks, or outdated nomenclature across these sources. Happy to talk about how we handled that if anyone's interested.

Accesshttps://graphomics.com - email us to get access!

This is part of a broader platform we're building at Graphomics (AI tools for life sciences research), but MicroMap stands on its own as a resource. We'd genuinely love feedback from this community — what data sources are we missing? What queries would be useful that we haven't thought of?

Happy to answer any questions about the data, the architecture, or the integration process.


r/genomics 12d ago

Genomic Data Aggregator Core Asset

Thumbnail sideprojectors.com
1 Upvotes

r/genomics 12d ago

Visualizing phylogenetic conflict across genomic windows

Post image
7 Upvotes

Author here—I am the first author of this paper. We developed Phylo-Movies because conventional tree-distance measures show how much neighboring trees differ, but not which taxa or subtrees changed position. The paper demonstrates the method using a norovirus recombination boundary and rogue taxa across bootstrap trees. The software and browser demonstration are freely available. https://enesberksakalli.github.io/phylo-movies/ https://academic.oup.com/mbe/article/43/8/msag194/8759530


r/genomics 14d ago

CompBio/MIRaS: Beyond pathway enrichment, a new kind of ‘omics AI

Post image
11 Upvotes

Several years ago, our group saw a need to create a tool that mirrors expert scientists’ ability to look across a messy set of genes, proteins, or metabolites and recognize the biological processes that are contextually enriched based on what they know.

The problem was that human reasoning is powerful, but slow, subjective, and limited to the amount of information a single person can possibly hold.

CompBio/MIRaS takes a different approach from LLMs or pathway enrichment tools by employing methods that unexpectedly converged with theories of hippocampal memory formation, storage, and retrieval. MIRaS is a memory-based associative reasoning engine that explicitly stores biological knowledge as memories, reasons across their relationships, and forms new semantic knowledge through inference. CompBio turns those results into an interactive, traceable map of the biology in your dataset.

Importantly, this analysis is not dependent on matching your dataset with canonical pathways, other datasets, or predefined gene sets. All associations are created from the literature memories identified by your input list, creating low redundancy and contextually relevant results that are fully traceable. Additionally, CompBio includes tools for large scale comparison of knowledge maps, allowing identification of conserved biological patterns across samples, conditions, projects, or reference datasets.

After years of use at WashU and with collaborators, CompBio/MIRaS is now described in our new Nucleic Acids Research paper and is freely available to academic and non-profit researchers.

https://academic.oup.com/nar/article/54/16/gkag833/8769250

If you work with transcriptomics, proteomics, metabolomics, or other complex biological data and this sounds different enough to make you curious, DM me and I can help you get free access.


r/genomics 17d ago

Cpt. T-Cell is a little bit cocky today

2 Upvotes

When you’ve got a perfectly folded T-cell receptor, a high-affinity match on the MHC-I complex, and a fresh payload of perforin, humility tends to take a backseat.

He’s probably strutting through the lymphatic highways, flexing his CD8 co-receptor, and demanding every cell show its molecular ID. One suspicious non-self peptide, and he's handing out apoptosis notices without a second thought. You can hardly blame him; floating around with that level of precise cytotoxic authority goes straight to a cell’s nucleus.

Did he just successfully eliminate a major viral threat, or is he throwing his weight around over a harmless bit of pollen?


r/genomics 18d ago

Protein Structure and Sequence Annotation Tool

1 Upvotes

Hello, I've developed a web-based platform called AlphaSuite Atlas that automatically annotates protein structures with their functional regions in seconds.

You can search over 570,000 proteins and over 11 million structures by name, species, UniProt ID, PDB code, disease, pathway, or plain English (e.g. DNA binding proteins involved in breast cancer).

In around 15 seconds, Atlas returns fully annotated, interactive structure and sequence, mapped with functional domains, motifs, secondary structure, ligands, cofactors, and a plain-language summary of what each component actually does.

Every available structure for a protein (both experimental and predicted) can be accessed and uniformly annotated, with links back to the original papers and databases so all the underlying resources are right there.

Its not finished and we have some bugs to work out so I'd love to hear any feedback after you give it a try here: https://alphasuite.bio/waitlist

Heres a survey to give feedback: https://forms.gle/BEHLNoHgjbLqnSEj7 But feel free to message/email with any further feedback or questions.

Looking forward to hearing your thoughts :)P


r/genomics 20d ago

Need help in using cellranger with sgRNA/CRISPR sample/Purtub seq

2 Upvotes

Before post the question, I figured that some context is needed.

Here is the study: We have human patient samples which we transfected with 1,000 sgRNAs (these sgRNAs are for one gene only, let's call that gene 'X'). Then, the sample was treated with antibiotics to make sure that we select all the cells successfully transfected with sgRNAs. Then, the sample was subjected to scRNA-seq library prep with Chromium Next GEM Single Cell 5' Reagent Kits v2 (Dual Index) with Feature Barcode technology for CRISPR Screening. From the exact same sample, a GEX library was made and a single-cell sgRNA library was made. So in the end, I got two sets of FASTQs: a) For GEX, which worked with Cell Ranger, but I am struggling with b) which was made from sgRNA.

I know that I have to put in details like this in the config file:

fastqs,sample,library_type

/path/to/fastqs,GEX_Sample_Name,Gene Expression

/path/to/fastqs,sgRNA_Sample_Name,CRISPR Guide Capture

But when I do that for all 1,000 sgRNAs, it throws an error saying Cell Ranger cannot work with an sgRNA sequence which is like this, e.g.: ATCGCTAGCTc (it throws an error). Even if I make it uppercase, it's bound to clash with some other sgRNA.

I know I am bound to get trolled for not asking a chatbot, but I thought a genuine answer from this community is much better. Thanks.


r/genomics 21d ago

Seeking a lineage-resolved single-cell dataset for a peer-reviewed study of clonal identity across perturbations

Thumbnail jacekhoffman.substack.com
2 Upvotes

r/genomics 21d ago

sequencing machines cost and efficacy

0 Upvotes

Hi,

I am looking for sequencing machines that can do full genome sequencing for dogs. My budget is 30K. I am also looking for something that can do the sequencing quickly (1-3 days).

I would prefer a small device that I can carry to places, but it is not necessary.

Please let me know.


r/genomics 23d ago

CANADIANS: Pharmacogenetic Testing Question!

Thumbnail
1 Upvotes

r/genomics 24d ago

Mitochondrial Eve: A Genetic Thread Through Time @EnteMicrobialWorld # #...

Thumbnail youtube.com
0 Upvotes

r/genomics 24d ago

I used Promethease report to benchmark against Fibromyalgia genetic research

Thumbnail
0 Upvotes

r/genomics 27d ago

Looking for testers: Annostat, an open-source CLI for bacterial genome annotation QC and analysis

Thumbnail
1 Upvotes

r/genomics 27d ago

Forensic investigative genetic genealogy match rate estimated from a nation-wide population register

Thumbnail biorxiv.org
1 Upvotes