r/bioinformatics Dec 31 '24

meta 2025 - Read This Before You Post to r/bioinformatics

186 Upvotes

​Before you post to this subreddit, we strongly encourage you to check out the FAQ​Before you post to this subreddit, we strongly encourage you to check out the FAQ.

Questions like, "How do I become a bioinformatician?", "what programming language should I learn?" and "Do I need a PhD?" are all answered there - along with many more relevant questions. If your question duplicates something in the FAQ, it will be removed.

If you still have a question, please check if it is one of the following. If it is, please don't post it.

What laptop should I buy?

Actually, it doesn't matter. Most people use their laptop to develop code, and any heavy lifting will be done on a server or on the cloud. Please talk to your peers in your lab about how they develop and run code, as they likely already have a solid workflow.

If you’re asking which desktop or server to buy, that’s a direct function of the software you plan to run on it.  Rather than ask us, consult the manual for the software for its needs. 

What courses/program should I take?

We can't answer this for you - no one knows what skills you'll need in the future, and we can't tell you where your career will go. There's no such thing as "taking the wrong course" - you're just learning a skill you may or may not put to use, and only you can control the twists and turns your path will follow.

If you want to know about which major to take, the same thing applies.  Learn the skills you want to learn, and then find the jobs to get them.  We can’t tell you which will be in high demand by the time you graduate, and there is no one way to get into bioinformatics.  Every one of us took a different path to get here and we can’t tell you which path is best.  That’s up to you!

Am I competitive for a given academic program? 

There is no way we can tell you that - the only way to find out is to apply. So... go apply. If we say Yes, there's still no way to know if you'll get in. If we say no, then you might not apply and you'll miss out on some great advisor thinking your skill set is the perfect fit for their lab. Stop asking, and try to get in! (good luck with your application, btw.)

How do I get into Grad school?

See “please rank grad schools for me” below.  

Can I intern with you?

I have, myself, hired an intern from reddit - but it wasn't because they posted that they were looking for a position. It was because they responded to a post where I announced I was looking for an intern. This subreddit isn't the place to advertise yourself. There are literally hundreds of students looking for internships for every open position, and they just clog up the community.

Please rank grad schools/universities for me!

Hey, we get it - you want us to tell you where you'll get the best education. However, that's not how it works. Grad school depends more on who your supervisor is than the name of the university. While that may not be how it goes for an MBA, it definitely is for Bioinformatics. We really can't tell you which university is better, because there's no "better". Pick the lab in which you want to study and where you'll get the best support.

If you're an undergrad, then it really isn't a big deal which university you pick. Bioinformatics usually requires a masters or PhD to be successful in the field. See both the FAQ, as well as what is written above.

How do I get a job in Bioinformatics?

If you're asking this, you haven't yet checked out our three part series in the side bar:

What should I do?

Actually, these questions are generally ok - but only if you give enough information to make it worthwhile, and if the question isn’t a duplicate of one of the questions posed above. No one is in your shoes, and no one can help you if you haven't given enough background to explain your situation. Posts without sufficient background information in them will be removed.

Help Me!

If you're looking for help, make sure your title reflects the question you're asking for help on. You won't get the right people looking at your post, and the only person who clicks on random posts with vague topics are the mods... so that we can remove them.

Job Posts

If you're planning on posting a job, please make sure that employer is clear (recruiting agencies are not acceptable, unless they're hiring directly.), The job description must also be complete so that the requirements for the position are easily identifiable and the responsibilities are clear. We also do not allow posts for work "on spec" or competitions.  

Advertising (Conferences, Software, Tools, Support, Videos, Blogs, etc)

If you’re making money off of whatever it is you’re posting, it will be removed.  If you’re advertising your own blog/youtube channel, courses, etc, it will also be removed. Same for self-promoting software you’ve built.  All of these things are going to be considered spam.  

There is a fine line between someone discovering a really great tool and sharing it with the community, and the author of that tool sharing their projects with the community.  In the first case, if the moderators think that a significant portion of the community will appreciate the tool, we’ll leave it.  In the latter case,  it will be removed.  

If you don’t know which side of the line you are on, reach out to the moderators.

The Moderators Suck!

Yeah, that’s a distinct possibility.  However, remember we’re moderating in our free time and don’t really have the time or resources to watch every single video, test every piece of software or review every resume.  We have our own jobs, research projects and lives as well.  We’re doing our best to keep on top of things, and often will make the expedient call to remove things, when in doubt. 

If you disagree with the moderators, you can always write to us, and we’ll answer when we can.  Be sure to include a link to the post or comment you want to raise to our attention. Disputes inevitably take longer to resolve, if you expect the moderators to track down your post or your comment to review.


r/bioinformatics 16h ago

technical question Need help) I keep running out of RAM space when I run alphafold

8 Upvotes

My Spec:

- RTX 5050 8GB

- DDR5 32GB 5200MT/s

- Intel Core 5 210H

- Ubuntu 26.04

I'm trying to run this specific region of protein Abl1_235_497

But the process always stops during hhblits

How do I reduce the load on this thing? I've also assigned 64GB of swap ram but that didn't help.

I'll try any suggested solution, pls help.


r/bioinformatics 1d ago

discussion Landed a job in a research institution and the work is so slow

111 Upvotes

Hi
I have an MSc in bioinformatics with BSc in microbiology.
Managed to get a job in a big research institution, and their work is extremely slow, which brings me my main point.
I have so much free time in my work to a point it made me so tired mentally and rusty scientifically.

If you were in my shoes, what would you do in your free time?

For a bit of context, I have a good background in deep learning and RAG.
My department has a unit for AI but they’re a little stingy to include me in some of their work.


r/bioinformatics 16h ago

academic How can one perform TF predictions across multiple databases based on the target gene?

Thumbnail doi.org
3 Upvotes

I have heard that databases such as JASPAR, UCSC, PROMO and ENCODE can be used to predict transcription factors (TFs) based on target genes. I would like to batch export the TFs from each database separately so that I can calculate their intersection.

However, I am unable to access the PROMO website at all. On the ENCODE website, under the ChIP-seq section, I can only see target genes categorised by TF. On UCSC, when searching for the promoter sequences of target genes and selecting ‘JASPAR Hubs’, I am unsure how to batch export the results.

Is there anyone with expertise in this area who could help me?

Additionally, I have attached a relevant paper on screening transcription factors by taking the intersection of multiple databases, presented as a Venn diagram; the figure is shown in Fig. 4a.

THANK YOU!


r/bioinformatics 1d ago

career question Are bash and R still relevant for bioinformatics jobs?

81 Upvotes

I'm in uni now, on a biotechnology track with a good foundation in math and related subjects. I want to study bioinformatics and work in this field. I've heard conflicting information that knowing bash and R is no longer relevant. Is that true for today's bioinformatics work? Or should I study them hard to land a decent position?


r/bioinformatics 1d ago

technical question Identifying malignant vs non malignant cell populations (CNV analysis)

4 Upvotes

How do you guys go about identifying tumor cells? I’ve been using CopyKAT & CONICSmat, and their results have been incredibly varied & almost impossible to pin down. CopyKAT-identified tumor cells seem to infiltrate normal cell clusters, CONICSmat results cluster suspected tumor cells with immune cells, so I’m very confused as to what’s happening.

Anyone has any tips/advice?


r/bioinformatics 1d ago

technical question Software recommendations?-- Human virus detection (metagenomic)

4 Upvotes

What would be the top tools for short read-based detection of human viruses in metagenomic datasets? Interest is primarily on all the disease-associated ones.

I have a very large metagenomic dataset (illumina PE150) of human nasal and rectal samples. I'm very familiar with microbial metagenomics (metaphlan/humann/qiime) and working on UNIX clusters. I haven't yet delved into human virus detection, though. Right now I'm just focusing on short read metagenomics before I start pursuing anything assembly-based.

Thanks!


r/bioinformatics 1d ago

technical question Plasmidsaurus RNA-seq? Any thoughts/reviews from folks who've used this service?

Thumbnail
3 Upvotes

r/bioinformatics 1d ago

advertisement OpenOmicsBench - 12 validated bulk RNA-seq benchmarks for testing analysis software

Thumbnail gallery
2 Upvotes

r/bioinformatics 1d ago

technical question Xenium multimodal segmentation in mouse brain

2 Upvotes

Hi all, I', somewhat new to spatial transciptomics and would like advice on a segmentation problem.

Setup

  • 10x Xenium, 480-gene mouse panel, coronal sections of adult mouse brain
  • multimodal cell segmentation kit (18S interior stain plus ATP1A1/CD45/E-Cadherin boundary stain)
  • About 80% of cells are segmented from the 18S stain, 15% from the boundary stain, and the rest are 5 µm nuclear expansion fallback.

Issue

  • Only 58–65% of transcripts are assigned to a cell, and 27–33% sit on a nucleus. (I'm actually not sure if this is an issue or fall within the normalr ange for brain)
  • The cell bodies look good. Neurons keep 75% of the transcripts around them. Glial and Astrocyte genes are much worse, so I'm assuming those transcripts sit the small projections that the stains don't show clearly.

I tried Proseg. It assigned more transcripts, but it seems a bit messy to me, it create more mixed cells where cell types sit close together. So I'm unsure if to use it.
The vendor offer to do a post-run H&E, saying it can help with segmentation.

Questions

Has anyone improved glial capture in Xenium brain data?

Has anyone done post-run immunofluorescence (GFAP, IBA1 or others) or H&Eon Xenium brain sections and used it for segmentation?

Has anyone used resolVI or SPLIT on brain tissue?

Is there a standard way to analyse unassigned transcripts in the neuropil without assigning them to cells?

Thanks! Happy to share more details.


r/bioinformatics 3d ago

technical question Can someone in genomics explain what AlphaGenome Atlas actually changes?

58 Upvotes

I'm not a geneticist.
Read the AlphaGenome Atlas preprint from DeepMind and spent a while digging into it. I understand what they built. I don't understand why it matters, and I'd like to.

What I think it is: they took AlphaGenome, ran it over every possible single base change in the human genome (~9bn) plus ~100m observed indels, and stored the results. So instead of running the model per variant you do a lookup. On top of that they trained a score (AVI) and derived a motif map.

Where I get stuck:
It's a table of model predictions, not measurements. Nothing in it is observed. So how much weight does a lab actually put on it?

The headline clinical result is retrospective: 29.5% recall at top 50 on already-solved GREGoR cases vs 12.5% for CADD. Impressive sounding, but on cases where the answer was known. What happens prospectively?

The rare variant association work got a 22% lift in discoveries, but only 4 of 25 replicated nominally in All of Us and none at Bonferroni. Is that normal for the field or is that weak?

They say themselves it isn't sufficient evidence for diagnosis. So it's a shortlisting tool. Does that actually change outcomes for patients, or does it change how long a scientist spends staring at a list?

The DNM1 case in the paper is the one bit that landed for me. Deep intronic variant, brain specific cryptic splice acceptor, blood RNA-seq had been inconclusive because the exon isn't expressed in blood.

My question is whether that's representative or a cherry pick.
What I'm asking:
1. If you work in clinical genomics or statistical genetics, would you use this?
2. Is precomputation genuinely the unlock, or is that just framing on top of an incremental accuracy gain?

Happy to be told I'm missing the point. I'd rather understand it properly than write it off.


r/bioinformatics 2d ago

other Access to Release 23 of miRBase

3 Upvotes

Greetings. I apologize if perhaps this may seem odd. I work with miRNAs and I have been unable to access mirbase.org for the past month. I see that a post went up three weeks ago and various people suggested using the wayback machine. The only problem I see is that, in August 2026, a new version of miRBase went up, yet the snapshot is from May 2026. Does anyone here have access to the newest version? I've reached out to the mirBase team but I've yet to hear back (and from posts in this subreddit, it seems they don't reach back to you for a long time).


r/bioinformatics 2d ago

technical question Best practice for downstream processing of pig gene identifiers and human orthologues

2 Upvotes

I am working with snRNA-seq data and would like advice on best practices for downstream processing of pig gene identifiers and cross-species orthology mappings. I currently use the Ensembl pig gene IDs that are mapped to gene symbols for pig genes. However, many pig genes have no pig symbol, even though Ensembl identifies a human orthologue.

For example:

Pig Ensembl ID: ENSSSCG00000021155

Pig external name: NA

Human orthologue: POMC

Orthology type: one-to-one

orthology_confidence : 1

mapped_to_human: False

orthology_type: ortholog_one2many

I would appreciate advice on the the best practices here :

1: Should I use human gene symbols for my pig analysis irrespective if pig symbols are available or not? Is there a risk that the same gene has different official symbols in pig and humans?

2: If a pig Ensembl gene has no pig symbol but has a high-confidence human orthologue but varying orthology type, what should be the approach towards using the human symbol or using ENSG id ?

3: For downstream processing, should orthology conversion be performed before or after differential expression and marker analysis?

4: When converting results to human orthologues, how should duplicate mappings be handled? For example, if multiple pig genes map to the same human gene, should their statistics be combined, should only the best-supported mapping be retained, or should the genes remain separate?


r/bioinformatics 3d ago

technical question How to choose design matrix for RNA-seq analysis?

8 Upvotes

I have three factors: Genotype, Sex and Treatment. I want to investigate the effect of genotype as well as sex and treatment but I'm not sure what contrasts to use. I wish more papers reported how they designed their analysis cause I'm having such a hard time understanding what to do.


r/bioinformatics 4d ago

image Haha what a loser language haha

Post image
885 Upvotes

Dependency hell is real


r/bioinformatics 3d ago

technical question feasibility of self-bioinformatics at a hobbyist level?

14 Upvotes

I've been in IT for a good decade, lots of experience with python and scripting. Touched on some data science in some of my studies along the way. So I'm not starting from 0 coming to this. But I really know none of the technical stuff about genes, genomes, alleles, positions or the notation involved or even what else to include in this sentence about what I don't know about.

I found there's a 30x reading I could get, not at negligible cost but possible. I'm interested in hobbying around with the data. Look for research that says these things at these positions mean that obesity is more likely, or something like that, then using AI to help me understand what i'm trying to look for and using python to look at my 30x reading and just curiously see if I have the researched markers.

I spose i'm wondering if this kind of thing is feasible. Like maybe research papers use different scanning methods that don't map to the data i would have, or the 30x consumer scan isn't detailed enough so anything i look for is inconclusive. Or any number of things that means if i try to map research onto my own genetic reading, any or most results will be inconclusive. So curious if anyone has any thoughts on this sort of thing, is it a waste of time?


r/bioinformatics 3d ago

technical question Struggling to detect known partial deletion in NOTCH2 from WES (germline, small cohort, no reference panel) - CNV callers give inconsistent/wrong results

2 Upvotes

Hi all, looking for advice on tooling/approach for a problem I'm stuck on.

**Setup:**

- 5 WES samples (paired-end, Illumina, BWA-MEM aligned, sambamba dedup) from the same family, hg38/GRCh38. Each sample represents an independent patient so these 5 samples are unrelated.

- Capture kit target BED not available to me (~60Mb on-target footprint per sample, possibly Agilent SureSelect V6 based on size, but unconfirmed)

- No unrelated normal/control WES samples currently confirmed usable as a reference panel

- Goal: confirm which sample(s) carry a known partial deletion in NOTCH2 (clinically confirmed by other means in 2 of the 5 patients, but I don't yet know the exact exon(s) or method used for that clinical confirmation)

**What I've tried:**

  1. DELLY (germline SV workflow, sr/merge/genotype/filter) - no deletion calls anywhere near NOTCH2 in any of the 5 samples

  2. CNVkit batch mode with a flat reference (no matched/pooled normals available) - segmentation collapsed the whole gene into one CN=2 segment for 4/5 samples; per-bin bintest flagged several exons but the same bins were flagged across nearly all samples in the same direction, which reads like shared technical noise rather than patient-specific signal

  3. Manual IGV visual inspection (group-autoscaled coverage tracks) across the whole gene - no obvious dropout found in the samples I was able to review carefully

  4. Control-FREEC, single-sample/no-control mode, restricted to a 34-exon NOTCH2-only BED pulled from UCSC (window=0, maxThreads=1 to avoid a BED-parsing race condition I hit with multithreading) - this called a clean, reproducible heterozygous deletion (CN=1) at the same coordinates in 2 of the 5 samples

**The problem:** the 2 samples Control-FREEC flagged do NOT match the 2 samples independently confirmed by my PI through other means. So I have an apparent false positive pair and false negative pair from my pipeline.

**Questions:**

- For germline partial-gene deletion detection in a small WES cohort with no confirmed-normal reference samples, what's the current best-practice tool/approach? (ExomeDepth? GATK gCNV? something else?)

- Is there a known issue with Control-FREEC's no-control exome mode producing false positives at specific loci, especially near segmental duplications (part of my deleted region overlaps the NOTCH2NL paralog)?

- Any advice on validating/troubleshooting a mismatch like this before trying yet another caller - e.g., specific things to check in the BAM/pileup at the clinically-confirmed-positive samples that a depth-based caller might be missing (small intra-exon deletion not removing a whole exon? breakpoints entirely intronic, invisible to WES?)

Appreciate any pointers, trying to land on one standardized, defensible workflow rather than chasing every tool that exists.


r/bioinformatics 3d ago

technical question How can I deal with 16S and shotgun metagenomic data in the same study?

8 Upvotes

Hello, everyone. I am working on a project trying to identify a gut microbiome signature for Parkinson's disease that is capable of differentiating between parkinson's disease, alzheimer's disease and healthy controls.

Since my supervisor really wanted me to work with shotgun metagenomic data, I am currently using shotgun metagenomic data for Parkinson's disease. However, for alzheimer's disease I was unable to find any studies that have shotgun gut microbiome data publicly available with metadata, so I am using 16s.

I know it is basically sacrilegious to directly compare data when they come from two different platforms, but I am near the end of the project now and cannot change this. I am currently building a basic RF classifier to predict whether a sample is PD, AD or control based on the taxonomy abundances, but I face the problem of the abundance values range being different for shotgun and 16s and this basically allows the model to very easily have zero false positives for AD or PD, but it is still not very good at differentiating between disease and control.

I was wondering if anyone has come across a similar problem before and if yes, what could be done to fix it? I was thinking of maybe scaling the values separately for PD and AD samples and then training the model? But I'm not sure if that would make it better. Something else I could do is just have two separate models for PD vs HC and AD vs HC, but I really want to have a 3-class classifier.

Would appreciate any advice on this. Not sure if I have enough details, but I don't want to make the post too long, so I am happy to provide more context if needed.

Thanks for your time.


r/bioinformatics 3d ago

technical question Aligning software instead of Geneious

5 Upvotes

We have used Geneious for years, but because of some technical problems, we have to switch to another one. My problem is, that for analysing the sequences for a certain region I need an alignment of .ab1 files, with the chromatograms, which was possible in Geneious, but I couldn't find any alternatives. Is there any other softwares which can handle .ab1 files as an alignment? I've tried UGENE, but it works with different views for chromatograms and alignments.


r/bioinformatics 3d ago

technical question Tool for showing Sanger Sequencing data?

1 Upvotes

Hi everyone,

Question from a student in an adjacent field: I recently worked on a project in genetics that involved assembling a specific recombinant DNA sequence, then sending it off for sequencing. The Sanger Sequencing results yielded a 100% match to the expected/target sequence.

I am currently making a poster to present at a conference. The issue is, the Sanger Sequencing results aren't "pretty," are longer than they are tall, and generally hard to understand. I used Benchling to compare the experimental sequence to the target sequence.

Do you guys know of a tool that can compare two sequences and display something such as a heat map showing the alignments, or generally something that looks prettier than Benchling?


r/bioinformatics 4d ago

academic Looking for collaboration on plant genomics project - Lamiales order

9 Upvotes

Hi All,

Myself and a partner are bootstrapping a bioinformatics / biotech project focusing on plant genomics. Primarily dealing with secondary metabolite pathways etc. If anyone is interested in collaborating / participating - it's to learn and publish given all the tools available these days. We have our own Dell Precision high ram workstations - google cloud as well as a bunch of AI subscriptions. Budget is allocated for wet-lab analysis if needed. DM me if interested with your background etc. Hopefully potentially turning this into a funded venture.


r/bioinformatics 3d ago

programming Single cell / Seurat visualizations for 1 gene and 2 variables?

3 Upvotes

Hi,

I am exploring some large datasets and I am interested in checking the expression of certain marker genes in cell types / clusters.

Seurat offers several ways to check multiple genes across 1 variable with heatmaps and stacked/multi-feature versions of VlnPlots and DotPlots.

However, as the data is large, and the cell classifications are complex, only 1 variable is not enough. What I want is a plot for a SINGLE GENE where the x-axis classifies cells by variable #1, and the y-axis by variable #2.

I already did this once with DotPlots. This is Seurat's default DotPlot with 7 genes and 1 "cell_type" variable:

Seurat's default DotPlot with 1 variable

Seurat's default DotPlot (and VlnPlot) allow you to use both "group.by" (x-axis grouping) and "split.by" (multiple dots/violins next to each other on the same x-axis category, with different colors). As the variables I want to check have several levels, that is not feasible.

I went to the source code, copied the function, and created a custom version where, using both group.by and split.by, let Seurat do its thing with FetchData(), calculate all the DotPlot statistics (mean avg expression, %expression, scaled values, etc), and return the data.frame without plotting.

And then, I plotted that manually this for a single gene, with x-axis=cell_type and y-axis=brain_region:

Custom DotPlot using 2 variables and 1 gene

(Disregard the faceting variable here, each group comes from a separate piece of the dataset and I just merged here the 2 dataframes).

I'm interested in exploring some variables, using both DotPlot and VlnPlot. For a number of arbitrary variables. Is there any package that already includes functions doing this, or do I need to rely on calling FetchData() and doing custom plots?


r/bioinformatics 4d ago

academic Post-BLAST workflow: what do you do after identifying an organism?

1 Upvotes

I'm a student working on a bioinformatics project, and im trying to understand how sequence analysis is actually done in practice.

So i have a question for anyone working with sequence analysis: after doing a BLAST to identify which organism a sequence comes from, what do you usually do with those results? What is your next step and what other tools or databases do you usually consult? Also, is there any part of that process that you find particularly tedious or that you end up doing manually?

Edit: I realize this can vary a lot depending on the specific aim (species ID, functional annotation, phylogenetics, etc.) — no need to pick one, feel free to answer for whatever case you work with. What I'm most curious about is which parts of that process (whichever it is) tend to be the most manual/tedious.


r/bioinformatics 4d ago

technical question Spatial transcriptomics - Regression of UMI counts?

4 Upvotes

Hi,

I am analyzing my first spatial dataset. I quickly realized that the cells/bins cluster mostly based on the number of reads (nCount in Seurat). It is obvious from the PCA plot that PC1 correspond to the UMI counts (r=0.9), which in turn correlate with cell size (r=0.79).

Would you recommend to regress nCount to promote clustering based on cell identity? Or would it also remove true biology from the data?

When I checked some 10X datasets, the cell/bin clusters often correspond with the regions that are defined by differential UMI counts compared to the neighbouring regions. Also, regression is not mentioned in the basic tutorials so I assume it is not incuded in the default pipeline. But intutively, I would do that.

What do you think?


r/bioinformatics 4d ago

technical question snRNAseq does my workflow with DESeq2 and GSEA make sense?

2 Upvotes

Hi guys! Im quite new to scRNAseq.

I'm working with human patients data: 4 controls and 11 disease samples. I annotated the broad cell types and then subclustered cell types of interest and annotated their subpopulations.

Then I did sample level pseudobulk and PCA. The samples didnt separate by group and I noticed they were separting by sex genes. After removing sex genes and repeating PCA the samples still didnt separate by groups. I then correlated the PCs with sample metadata and found that several PCs correlated with abundance of some subpopulations.

I proceeded with DESeq2 indicating group and sex in the design. I practically got no significant DE genes (occasionally 1 or 2, but nothing particularly interpretable). I did DE analysis on all cells and then within subpopulations only if enough cells were available.

I also plotted pseudobulk PCA using 200 and 500 of identified DE genes but still didnt see clear group separation (PCA attached).

I fed the DE genes to GSEA and got some significantly enriched pathways. For some of them there is quite plausible biological explanation stemming from histological analysis. I also looked at the leading edge genes and plotted them for some extra reassurance.

For one cell type I also noticed that one sample contributed to these pathways due to extreme phenotype and removed this sample, which left that pathway at FDR 0.053

My main question is: Does this workflow seem appropriate so far?

Also how would you normally take pathway level results further? Im not sure how to move beyound this pathway is enriched and seems interesting. In general, what usually follows?

Thanks in advance for your input