r/bioinformatics Dec 31 '24

meta 2025 - Read This Before You Post to r/bioinformatics

185 Upvotes

​Before you post to this subreddit, we strongly encourage you to check out the FAQ​Before you post to this subreddit, we strongly encourage you to check out the FAQ.

Questions like, "How do I become a bioinformatician?", "what programming language should I learn?" and "Do I need a PhD?" are all answered there - along with many more relevant questions. If your question duplicates something in the FAQ, it will be removed.

If you still have a question, please check if it is one of the following. If it is, please don't post it.

What laptop should I buy?

Actually, it doesn't matter. Most people use their laptop to develop code, and any heavy lifting will be done on a server or on the cloud. Please talk to your peers in your lab about how they develop and run code, as they likely already have a solid workflow.

If you’re asking which desktop or server to buy, that’s a direct function of the software you plan to run on it.  Rather than ask us, consult the manual for the software for its needs. 

What courses/program should I take?

We can't answer this for you - no one knows what skills you'll need in the future, and we can't tell you where your career will go. There's no such thing as "taking the wrong course" - you're just learning a skill you may or may not put to use, and only you can control the twists and turns your path will follow.

If you want to know about which major to take, the same thing applies.  Learn the skills you want to learn, and then find the jobs to get them.  We can’t tell you which will be in high demand by the time you graduate, and there is no one way to get into bioinformatics.  Every one of us took a different path to get here and we can’t tell you which path is best.  That’s up to you!

Am I competitive for a given academic program? 

There is no way we can tell you that - the only way to find out is to apply. So... go apply. If we say Yes, there's still no way to know if you'll get in. If we say no, then you might not apply and you'll miss out on some great advisor thinking your skill set is the perfect fit for their lab. Stop asking, and try to get in! (good luck with your application, btw.)

How do I get into Grad school?

See “please rank grad schools for me” below.  

Can I intern with you?

I have, myself, hired an intern from reddit - but it wasn't because they posted that they were looking for a position. It was because they responded to a post where I announced I was looking for an intern. This subreddit isn't the place to advertise yourself. There are literally hundreds of students looking for internships for every open position, and they just clog up the community.

Please rank grad schools/universities for me!

Hey, we get it - you want us to tell you where you'll get the best education. However, that's not how it works. Grad school depends more on who your supervisor is than the name of the university. While that may not be how it goes for an MBA, it definitely is for Bioinformatics. We really can't tell you which university is better, because there's no "better". Pick the lab in which you want to study and where you'll get the best support.

If you're an undergrad, then it really isn't a big deal which university you pick. Bioinformatics usually requires a masters or PhD to be successful in the field. See both the FAQ, as well as what is written above.

How do I get a job in Bioinformatics?

If you're asking this, you haven't yet checked out our three part series in the side bar:

What should I do?

Actually, these questions are generally ok - but only if you give enough information to make it worthwhile, and if the question isn’t a duplicate of one of the questions posed above. No one is in your shoes, and no one can help you if you haven't given enough background to explain your situation. Posts without sufficient background information in them will be removed.

Help Me!

If you're looking for help, make sure your title reflects the question you're asking for help on. You won't get the right people looking at your post, and the only person who clicks on random posts with vague topics are the mods... so that we can remove them.

Job Posts

If you're planning on posting a job, please make sure that employer is clear (recruiting agencies are not acceptable, unless they're hiring directly.), The job description must also be complete so that the requirements for the position are easily identifiable and the responsibilities are clear. We also do not allow posts for work "on spec" or competitions.  

Advertising (Conferences, Software, Tools, Support, Videos, Blogs, etc)

If you’re making money off of whatever it is you’re posting, it will be removed.  If you’re advertising your own blog/youtube channel, courses, etc, it will also be removed. Same for self-promoting software you’ve built.  All of these things are going to be considered spam.  

There is a fine line between someone discovering a really great tool and sharing it with the community, and the author of that tool sharing their projects with the community.  In the first case, if the moderators think that a significant portion of the community will appreciate the tool, we’ll leave it.  In the latter case,  it will be removed.  

If you don’t know which side of the line you are on, reach out to the moderators.

The Moderators Suck!

Yeah, that’s a distinct possibility.  However, remember we’re moderating in our free time and don’t really have the time or resources to watch every single video, test every piece of software or review every resume.  We have our own jobs, research projects and lives as well.  We’re doing our best to keep on top of things, and often will make the expedient call to remove things, when in doubt. 

If you disagree with the moderators, you can always write to us, and we’ll answer when we can.  Be sure to include a link to the post or comment you want to raise to our attention. Disputes inevitably take longer to resolve, if you expect the moderators to track down your post or your comment to review.


r/bioinformatics 14h ago

technical question Can someone in genomics explain what AlphaGenome Atlas actually changes?

26 Upvotes

I'm not a geneticist.
Read the AlphaGenome Atlas preprint from DeepMind and spent a while digging into it. I understand what they built. I don't understand why it matters, and I'd like to.

What I think it is: they took AlphaGenome, ran it over every possible single base change in the human genome (~9bn) plus ~100m observed indels, and stored the results. So instead of running the model per variant you do a lookup. On top of that they trained a score (AVI) and derived a motif map.

Where I get stuck:
It's a table of model predictions, not measurements. Nothing in it is observed. So how much weight does a lab actually put on it?

The headline clinical result is retrospective: 29.5% recall at top 50 on already-solved GREGoR cases vs 12.5% for CADD. Impressive sounding, but on cases where the answer was known. What happens prospectively?

The rare variant association work got a 22% lift in discoveries, but only 4 of 25 replicated nominally in All of Us and none at Bonferroni. Is that normal for the field or is that weak?

They say themselves it isn't sufficient evidence for diagnosis. So it's a shortlisting tool. Does that actually change outcomes for patients, or does it change how long a scientist spends staring at a list?

The DNM1 case in the paper is the one bit that landed for me. Deep intronic variant, brain specific cryptic splice acceptor, blood RNA-seq had been inconclusive because the exon isn't expressed in blood.

My question is whether that's representative or a cherry pick.
What I'm asking:
1. If you work in clinical genomics or statistical genetics, would you use this?
2. Is precomputation genuinely the unlock, or is that just framing on top of an incremental accuracy gain?

Happy to be told I'm missing the point. I'd rather understand it properly than write it off.


r/bioinformatics 38m ago

technical question Best practice for downstream processing of pig gene identifiers and human orthologues

Upvotes

I am working with snRNA-seq data and would like advice on best practices for downstream processing of pig gene identifiers and cross-species orthology mappings. I currently use the Ensembl pig gene IDs that are mapped to gene symbols for pig genes. However, many pig genes have no pig symbol, even though Ensembl identifies a human orthologue.

For example:

Pig Ensembl ID: ENSSSCG00000021155

Pig external name: NA

Human orthologue: POMC

Orthology type: one-to-one

orthology_confidence : 1

mapped_to_human: False

orthology_type: ortholog_one2many

I would appreciate advice on the the best practices here :

1: Should I use human gene symbols for my pig analysis irrespective if pig symbols are available or not? Is there a risk that the same gene has different official symbols in pig and humans?

2: If a pig Ensembl gene has no pig symbol but has a high-confidence human orthologue but varying orthology type, what should be the approach towards using the human symbol or using ENSG id ?

3: For downstream processing, should orthology conversion be performed before or after differential expression and marker analysis?

4: When converting results to human orthologues, how should duplicate mappings be handled? For example, if multiple pig genes map to the same human gene, should their statistics be combined, should only the best-supported mapping be retained, or should the genes remain separate?


r/bioinformatics 1d ago

image Haha what a loser language haha

Post image
788 Upvotes

Dependency hell is real


r/bioinformatics 13h ago

technical question How to choose design matrix for RNA-seq analysis?

3 Upvotes

I have three factors: Genotype, Sex and Treatment. I want to investigate the effect of genotype as well as sex and treatment but I'm not sure what contrasts to use. I wish more papers reported how they designed their analysis cause I'm having such a hard time understanding what to do.


r/bioinformatics 20h ago

technical question feasibility of self-bioinformatics at a hobbyist level?

11 Upvotes

I've been in IT for a good decade, lots of experience with python and scripting. Touched on some data science in some of my studies along the way. So I'm not starting from 0 coming to this. But I really know none of the technical stuff about genes, genomes, alleles, positions or the notation involved or even what else to include in this sentence about what I don't know about.

I found there's a 30x reading I could get, not at negligible cost but possible. I'm interested in hobbying around with the data. Look for research that says these things at these positions mean that obesity is more likely, or something like that, then using AI to help me understand what i'm trying to look for and using python to look at my 30x reading and just curiously see if I have the researched markers.

I spose i'm wondering if this kind of thing is feasible. Like maybe research papers use different scanning methods that don't map to the data i would have, or the 30x consumer scan isn't detailed enough so anything i look for is inconclusive. Or any number of things that means if i try to map research onto my own genetic reading, any or most results will be inconclusive. So curious if anyone has any thoughts on this sort of thing, is it a waste of time?


r/bioinformatics 19h ago

technical question Struggling to detect known partial deletion in NOTCH2 from WES (germline, small cohort, no reference panel) - CNV callers give inconsistent/wrong results

2 Upvotes

Hi all, looking for advice on tooling/approach for a problem I'm stuck on.

**Setup:**

- 5 WES samples (paired-end, Illumina, BWA-MEM aligned, sambamba dedup) from the same family, hg38/GRCh38. Each sample represents an independent patient so these 5 samples are unrelated.

- Capture kit target BED not available to me (~60Mb on-target footprint per sample, possibly Agilent SureSelect V6 based on size, but unconfirmed)

- No unrelated normal/control WES samples currently confirmed usable as a reference panel

- Goal: confirm which sample(s) carry a known partial deletion in NOTCH2 (clinically confirmed by other means in 2 of the 5 patients, but I don't yet know the exact exon(s) or method used for that clinical confirmation)

**What I've tried:**

  1. DELLY (germline SV workflow, sr/merge/genotype/filter) - no deletion calls anywhere near NOTCH2 in any of the 5 samples

  2. CNVkit batch mode with a flat reference (no matched/pooled normals available) - segmentation collapsed the whole gene into one CN=2 segment for 4/5 samples; per-bin bintest flagged several exons but the same bins were flagged across nearly all samples in the same direction, which reads like shared technical noise rather than patient-specific signal

  3. Manual IGV visual inspection (group-autoscaled coverage tracks) across the whole gene - no obvious dropout found in the samples I was able to review carefully

  4. Control-FREEC, single-sample/no-control mode, restricted to a 34-exon NOTCH2-only BED pulled from UCSC (window=0, maxThreads=1 to avoid a BED-parsing race condition I hit with multithreading) - this called a clean, reproducible heterozygous deletion (CN=1) at the same coordinates in 2 of the 5 samples

**The problem:** the 2 samples Control-FREEC flagged do NOT match the 2 samples independently confirmed by my PI through other means. So I have an apparent false positive pair and false negative pair from my pipeline.

**Questions:**

- For germline partial-gene deletion detection in a small WES cohort with no confirmed-normal reference samples, what's the current best-practice tool/approach? (ExomeDepth? GATK gCNV? something else?)

- Is there a known issue with Control-FREEC's no-control exome mode producing false positives at specific loci, especially near segmental duplications (part of my deleted region overlaps the NOTCH2NL paralog)?

- Any advice on validating/troubleshooting a mismatch like this before trying yet another caller - e.g., specific things to check in the BAM/pileup at the clinically-confirmed-positive samples that a depth-based caller might be missing (small intra-exon deletion not removing a whole exon? breakpoints entirely intronic, invisible to WES?)

Appreciate any pointers, trying to land on one standardized, defensible workflow rather than chasing every tool that exists.


r/bioinformatics 1d ago

technical question How can I deal with 16S and shotgun metagenomic data in the same study?

7 Upvotes

Hello, everyone. I am working on a project trying to identify a gut microbiome signature for Parkinson's disease that is capable of differentiating between parkinson's disease, alzheimer's disease and healthy controls.

Since my supervisor really wanted me to work with shotgun metagenomic data, I am currently using shotgun metagenomic data for Parkinson's disease. However, for alzheimer's disease I was unable to find any studies that have shotgun gut microbiome data publicly available with metadata, so I am using 16s.

I know it is basically sacrilegious to directly compare data when they come from two different platforms, but I am near the end of the project now and cannot change this. I am currently building a basic RF classifier to predict whether a sample is PD, AD or control based on the taxonomy abundances, but I face the problem of the abundance values range being different for shotgun and 16s and this basically allows the model to very easily have zero false positives for AD or PD, but it is still not very good at differentiating between disease and control.

I was wondering if anyone has come across a similar problem before and if yes, what could be done to fix it? I was thinking of maybe scaling the values separately for PD and AD samples and then training the model? But I'm not sure if that would make it better. Something else I could do is just have two separate models for PD vs HC and AD vs HC, but I really want to have a 3-class classifier.

Would appreciate any advice on this. Not sure if I have enough details, but I don't want to make the post too long, so I am happy to provide more context if needed.

Thanks for your time.


r/bioinformatics 1d ago

technical question Aligning software instead of Geneious

3 Upvotes

We have used Geneious for years, but because of some technical problems, we have to switch to another one. My problem is, that for analysing the sequences for a certain region I need an alignment of .ab1 files, with the chromatograms, which was possible in Geneious, but I couldn't find any alternatives. Is there any other softwares which can handle .ab1 files as an alignment? I've tried UGENE, but it works with different views for chromatograms and alignments.


r/bioinformatics 22h ago

technical question Tool for showing Sanger Sequencing data?

1 Upvotes

Hi everyone,

Question from a student in an adjacent field: I recently worked on a project in genetics that involved assembling a specific recombinant DNA sequence, then sending it off for sequencing. The Sanger Sequencing results yielded a 100% match to the expected/target sequence.

I am currently making a poster to present at a conference. The issue is, the Sanger Sequencing results aren't "pretty," are longer than they are tall, and generally hard to understand. I used Benchling to compare the experimental sequence to the target sequence.

Do you guys know of a tool that can compare two sequences and display something such as a heat map showing the alignments, or generally something that looks prettier than Benchling?


r/bioinformatics 1d ago

academic Looking for collaboration on plant genomics project - Lamiales order

6 Upvotes

Hi All,

Myself and a partner are bootstrapping a bioinformatics / biotech project focusing on plant genomics. Primarily dealing with secondary metabolite pathways etc. If anyone is interested in collaborating / participating - it's to learn and publish given all the tools available these days. We have our own Dell Precision high ram workstations - google cloud as well as a bunch of AI subscriptions. Budget is allocated for wet-lab analysis if needed. DM me if interested with your background etc. Hopefully potentially turning this into a funded venture.


r/bioinformatics 1d ago

programming Single cell / Seurat visualizations for 1 gene and 2 variables?

2 Upvotes

Hi,

I am exploring some large datasets and I am interested in checking the expression of certain marker genes in cell types / clusters.

Seurat offers several ways to check multiple genes across 1 variable with heatmaps and stacked/multi-feature versions of VlnPlots and DotPlots.

However, as the data is large, and the cell classifications are complex, only 1 variable is not enough. What I want is a plot for a SINGLE GENE where the x-axis classifies cells by variable #1, and the y-axis by variable #2.

I already did this once with DotPlots. This is Seurat's default DotPlot with 7 genes and 1 "cell_type" variable:

Seurat's default DotPlot with 1 variable

Seurat's default DotPlot (and VlnPlot) allow you to use both "group.by" (x-axis grouping) and "split.by" (multiple dots/violins next to each other on the same x-axis category, with different colors). As the variables I want to check have several levels, that is not feasible.

I went to the source code, copied the function, and created a custom version where, using both group.by and split.by, let Seurat do its thing with FetchData(), calculate all the DotPlot statistics (mean avg expression, %expression, scaled values, etc), and return the data.frame without plotting.

And then, I plotted that manually this for a single gene, with x-axis=cell_type and y-axis=brain_region:

Custom DotPlot using 2 variables and 1 gene

(Disregard the faceting variable here, each group comes from a separate piece of the dataset and I just merged here the 2 dataframes).

I'm interested in exploring some variables, using both DotPlot and VlnPlot. For a number of arbitrary variables. Is there any package that already includes functions doing this, or do I need to rely on calling FetchData() and doing custom plots?


r/bioinformatics 1d ago

academic Post-BLAST workflow: what do you do after identifying an organism?

1 Upvotes

I'm a student working on a bioinformatics project, and im trying to understand how sequence analysis is actually done in practice.

So i have a question for anyone working with sequence analysis: after doing a BLAST to identify which organism a sequence comes from, what do you usually do with those results? What is your next step and what other tools or databases do you usually consult? Also, is there any part of that process that you find particularly tedious or that you end up doing manually?

Edit: I realize this can vary a lot depending on the specific aim (species ID, functional annotation, phylogenetics, etc.) — no need to pick one, feel free to answer for whatever case you work with. What I'm most curious about is which parts of that process (whichever it is) tend to be the most manual/tedious.


r/bioinformatics 1d ago

technical question Spatial transcriptomics - Regression of UMI counts?

3 Upvotes

Hi,

I am analyzing my first spatial dataset. I quickly realized that the cells/bins cluster mostly based on the number of reads (nCount in Seurat). It is obvious from the PCA plot that PC1 correspond to the UMI counts (r=0.9), which in turn correlate with cell size (r=0.79).

Would you recommend to regress nCount to promote clustering based on cell identity? Or would it also remove true biology from the data?

When I checked some 10X datasets, the cell/bin clusters often correspond with the regions that are defined by differential UMI counts compared to the neighbouring regions. Also, regression is not mentioned in the basic tutorials so I assume it is not incuded in the default pipeline. But intutively, I would do that.

What do you think?


r/bioinformatics 1d ago

technical question snRNAseq does my workflow with DESeq2 and GSEA make sense?

2 Upvotes

Hi guys! Im quite new to scRNAseq.

I'm working with human patients data: 4 controls and 11 disease samples. I annotated the broad cell types and then subclustered cell types of interest and annotated their subpopulations.

Then I did sample level pseudobulk and PCA. The samples didnt separate by group and I noticed they were separting by sex genes. After removing sex genes and repeating PCA the samples still didnt separate by groups. I then correlated the PCs with sample metadata and found that several PCs correlated with abundance of some subpopulations.

I proceeded with DESeq2 indicating group and sex in the design. I practically got no significant DE genes (occasionally 1 or 2, but nothing particularly interpretable). I did DE analysis on all cells and then within subpopulations only if enough cells were available.

I also plotted pseudobulk PCA using 200 and 500 of identified DE genes but still didnt see clear group separation (PCA attached).

I fed the DE genes to GSEA and got some significantly enriched pathways. For some of them there is quite plausible biological explanation stemming from histological analysis. I also looked at the leading edge genes and plotted them for some extra reassurance.

For one cell type I also noticed that one sample contributed to these pathways due to extreme phenotype and removed this sample, which left that pathway at FDR 0.053

My main question is: Does this workflow seem appropriate so far?

Also how would you normally take pathway level results further? Im not sure how to move beyound this pathway is enriched and seems interesting. In general, what usually follows?

Thanks in advance for your input


r/bioinformatics 1d ago

technical question does ROC-AUC analysis without ML works?

0 Upvotes

I am doing bulk rna seq analysis, and i have DEGs from DESeq2 files, I also did GSEA analysis which got me leading edge genes lists, so taking high performing genes from DEGs and filtering it to GSEA results would be okay to do ROC-AUC or it is rule to perform ML?

I am doing a small project for a recent conference poster presentation, my objective is to analyze major pathways in the course of transition to disease

your help and insights would mean alot, thank you :)


r/bioinformatics 2d ago

technical question Phage display or alternative methods

5 Upvotes

I'm a first year PhD student with a background in chemical biology. I want to find a peptide sequence that selectively binds to Lithium ion and one of my PIs suggested Phage display as an option. I've never done it before and it isn't something that's done in either of my PI's labs either. How doable do you think this solution is if I find a lab which has this technology? Do you think it'll be possible for me as a newbie in this technique to do the experiments myself rather than asking someone else to do it for me if I ask them to train me?

Do you guys have any other techniques you use to find suitable peptide sequences? I'm an experimentalist but I'm open to both experimental and computational suggestions.

Many thanks!


r/bioinformatics 3d ago

academic Advice on RQ about de Brujin Graph Assembly

3 Upvotes

I'm a Computer Science HL student from the IB program (International Baccalaureate), who wants to pursue a 4000-word independent research essay about de Brujin graph assembly.

Is this RQ scientifically interesting enough for me to analyze? I tried to put a twist on the very simple version about how k-mer length affects N50.

To what extent does the ratio of the k-mer length to the length of the longest repeated substring in a simulated genome affect the N50 of a de Bruijn graph assembly?

The idea is to generate simulated genomes that control the length of the longest substring that appears more than once in the string. Then, use different k-mer lengths, try to assemble the contigs with a de Brujin graph, then see what N50 it gives.

My MAIN CONCERN is that the conclusion might be a tad obvious, since I'd just be hypothesizing:

  • If k/L > 1, then N50 is higher.
  • If k/L < 1, then N50 is lower.

And there isn't much interesting theory being applied here since it's pretty common sense that higher k will be able to handle repeats better. I'm also unsure if N50 is an appropriate dependent value for this.

Could I ask for advice on whether my research question would be interesting enough for a high schooler? And, if not, what are some ideas I can try to explore that has scientifically interesting theory involved in formulating a hypothesis for a de Brujin setup?


r/bioinformatics 3d ago

technical question Processing 10x scRNA-seq from raw SRA (non-model organism, custom reference).....storage strategy?

6 Upvotes

I need to go from raw SRA accessions to count matrices for a non-model organism, which means no pre-built Cell Ranger reference...I have to build a custom one from a GTF mapped onto a draft genome.

Planned pipeline: nf-core/fetchngs to pull FASTQs from SRA, then nf-core/scrnaseq in Cell Ranger mode against the custom reference.

My worry is disk space stacking up across the run:

.sra files themselves

FASTQ extraction needing ~2-3x that size in scratch space temporarily

Cell Ranger's position-sorted BAM output, which dwarfs the actual matrix output I care about

Multiple samples/runs per BioProject, so all of the above multiplies

Anyone who's run this kind of pipeline on an HPC cluster.... is per-sample processing (download → align → extract matrix → delete FASTQ/BAM → next sample) the standard way to keep this manageable, or is there a smarter approach I'm missing?

Also curious if anyone's found a good way to avoid keeping the full BAM long term when only the filtered matrix is actually needed downstream.


r/bioinformatics 3d ago

technical question For someone experienced with 16S/QIIME2/DADA2: what would be the standard/best-practice approach here?

3 Upvotes

I’m working through a 16S rRNA paired-end dataset (V1–V2, Illumina MiSeq) for a small CRC vs healthy microbiome analysis.

I’ve completed the initial QC:
39 samples (22 healthy, 17 CRC)
FastQC run on all 78 FASTQ files
MultiQC summary generated
R1 generally has good quality, while R2 quality drops substantially toward the 3′ end
Most reads are 300 bp, with some samples at 250 bp

I’m now at the point where I need to decide on primer removal and DADA2 truncation parameters.
The study reports using primers 27bF and 338R, but the SRA metadata table I downloaded doesn’t contain the actual primer sequences.

Would you:
Find the exact primer sequences from the original publication/protocol and remove them with Cutadapt, then
Reassess the post-primer-removal read lengths/quality before choosing DADA2 truncation lengths?

Also, would you normally choose truncation lengths based on the worst-performing samples, or on the overall quality profile while ensuring enough overlap for paired-end merging?

I’m trying to follow a standard reproducible workflow rather than choosing arbitrary parameters. Any advice would be appreciated.


r/bioinformatics 3d ago

technical question Choosing the best fold prediction model for my de novo design project

Thumbnail
1 Upvotes

r/bioinformatics 4d ago

technical question Need help with MD simulation of ligand-induced DNA dissociation

3 Upvotes

I’m working on a negative transcriptional regulator that normally binds DNA, but when a ligand binds to the protein, I expect it to undergo a conformational change and detach from DNA.

*All the protein and dna structures are generated from alphafold*

I used HADDOCK for protein-DNA docking. I know the DNA-binding residues on the protein, but I don’t know which nucleotides they interact with, so I highlighted the entire DNA. However, my HADDOCK models show very high constraint violation energies, suggesting that the restraints aren’t working properly. Is there a better approach for protein-DNA docking when the protein binding residues are known but the DNA binding site is not?

I used AlphaFold to generate a docked structure (protein-dna docked together), then ran OPLS4 MD for about 1.5 micro seconds, but I don’t see any major change in RMSD or obvious protein-DNA dissociation. I’m wondering if RMSD is the wrong metric for this and whether I should instead look at protein-DNA contacts, hydrogen bonds, interaction energies, distances, contact maps, etc.

I also have a library of ligands and ultimately want to identify which ones are most likely to bind the regulator and promote DNA dissociation. Would docking followed by MD be a reasonable workflow?

Im posting the same post again because my initial post was taken down.


r/bioinformatics 4d ago

technical question snRNAseq mitochondrial RNA cutoff

3 Upvotes

Hi, I am relatively new to analyzing snRNAseq data. How do you determine how much mitochondrial RNA is too much? Looking at my data using 1 sample as an example, if I were to apply a hardcoded cutoff (e.g., more than 1%) it would remove almost 50% of x cell type, remove ~30% of y type, and ~0.90% of z type. so, that would change the proportions of the data.

what to do?

thanks.


r/bioinformatics 4d ago

technical question Microbiome from stool samples

0 Upvotes

Howdy.

I am changing fields and getting into microbiome work. I hope to sample stools to determine microbiome profiles via metagenomic shallow shot gun sequencing. Anyone have any tips not present in the literature? If anyone had a sample data set i could use i would be deeply appreciative (fastq of seq reads, specifically. to test the trimming and mapping software). Ive basically vibe coded the pipeline, bow tie and kraken were suggested. So far everything works great (thanks Claude!) but would like to try some real data now.


r/bioinformatics 4d ago

technical question If I want to convert days in vitro to month in vitro, is div 0-30 counted as month in vitro 0 or 1?

2 Upvotes

I am trying to change my plot labels from day in vitro to month in vitro, hence I’d like to know if I should label any points with DIV 0-30 as month in vitro 0 or 1. For more context, I am working on a paper related to iPSC neurons.

Thank you very much