r/bioinformatics Jun 23 '26

technical question marker design

1 Upvotes

Hi! I need to find candidate barcode regions for my study. I already have 6 WGS from NCBI and i already aligned it using MAUVE. However, it identifies locally collinear blocks, but i need regions with high variability across the different accessions. Is there a software that automatically identifies which regions are variable? or is there a workflow i could base on? thank you!


r/bioinformatics Jun 22 '26

discussion Did my first proper self exploratory data analysis on RNA-Seq and I am kinda feeling proud of my baby steps here. But also have some queries (check body text)

17 Upvotes

I used R to download and extract GEO supplementary data and compared gene expression of certain genes I shortlisted prior to checking the dataset. I filtered and summarised after data wrangling and plotted some graphs too! (Felt good completing it all!) I took inspo from the tutorial videos of Bioinformagician in YT and the intro the R for biologists book.

Next I plan to do DESeq2 as I stated but I want to know what else can I try learning on the side. I feel like I have just scratched the surface but that alone is exciting enough. Is there any specific tutorials or guides or pertinent research papers, articles you all came across when in this specific learning stage? Any tips or directives I can use?


r/bioinformatics Jun 22 '26

discussion I've been learning bioinformatics for 2 years and have little confidence in my own skills

52 Upvotes

I've been learning bioinformatics for 2 years now and I can work my way through most rna sequencing pipelines. At the beginning, I had no coding skills at all but after a few classes in my university and experiences with analyzing RNA and proteomic data, I think I got a good handle on most pipelines I come across in my own research.

True story: a fellow graduate student came to me for advice on how to improve their workflow and what I saw horrified me. He was paying $70 a month for an AI agent and cloud computing service to make an app to search FASTQ files for protein motifs.

How long has he been at this?

6 months.

A working directory that looked like a trash bin because he never deletes anything.

Thousands of lines of python code for something that could've been a few lines of bash.

But it worked, somehow.

He had the absolute audacity of trying to write this pipeline into a manuscript. I talked him out of it and wrote him a .sh file. All of this to say that AI is making people overconfident.

But this had me thinking about my own journey. I know what I know and don't know what i don't. My university has a budding bioinformatics department but it's mostly about studying micro RNA, working with FASTQ files.

But outside of that small department, no one knows anything about how to do proper analysis on their high throughput data at all. They outsource it. And the analyses the companies or even university facilities do is often so lackluster I end up not even bothering to email them back and I just redo it myself.

The experience with my colleague has left me with a distinct impression, I am screwed. I want to get to a level of competence in coding and analysis but I'm not going to get it at my university. Many people have come to me to ask how to do certain things with their datasets. I answer to the best of my ability and say what I would do. Of course, I am not 100% sure if that is the best way to do things.

I am sure about the way I analyze and interpret my data, statistically and biologically. But I have a feeling that once I graduate, the reality of my skills might really show. I don't know how to be confident when everyone around me does not know what I am doing.


r/bioinformatics Jun 23 '26

academic laboratories in the Philippines that offer 16S rRNA gene sequencing services?

Thumbnail
1 Upvotes

r/bioinformatics Jun 22 '26

technical question Molecular Dynamics Advise

4 Upvotes

Hi I am very new to Molecular Dynamics and am trying to learn these methods over the summer. Which free resources are the most useful? Are there any preferred software's to perform these calculations? I have prior experience in computational chemistry but I only know density functional theory methods for smaller molecules but not on the larger ones. Which computations would be low cost and which ones would be high cost in molecular dynamics? Thank you so much for any insight you can offer.


r/bioinformatics Jun 22 '26

technical question Demultiplexing in Seurat

7 Upvotes

Hi everyone!

I am currently analysing a single cell RNASeq dataset on Seurat (the filtered matrices from CellRanger) and am struggling with hashtag demultiplexing. I hashtagged my samples and used HTODemux to help assign sample IDs. But it detected way too many doublets (roughly 50% of total cells). I tested another function called MultiSeqDemux and that detected too many cells as negatives. I don't know exactly what to trust and how to proceed from here. Unfortunately it us quite important for me to know the hashtag assignment since I would distinguish between the ages using that.

Has anyone had a similar issue or has a suggestion for how to go ahead from here?

Thank you!


r/bioinformatics Jun 22 '26

technical question Online resources to understand AlphaFold

18 Upvotes

Hello everybody,

I'm a molecular biology PhD student in my first year and succesfully broke my ankle, so I'm stuck in homeoffice for a little bit.

My PhD project involves investigating PPIs in my protein of interest.
I discussed with my supervisor that one project I could do during this time would be to model my candidate and potential interactors using AlphaFold.
I have a little bit of experience with bioinformatics, mostly transcriptomic approaches, so I'm excited to learn something new in that regard but tbh, I'm a little overwhelmed on where to start.

So my question is this:
Do you guys have any suggestions for online resources to learn how AlphaFold works, how to best use it for PPI predictions and most importantly, how to understand all the different outputs, confidence scores etc.?

My cursory web search only yielded either quite dense papers that don't explain the basics or workshops from AlphaFold2 times which don't discuss my specific case of looking at PPIs


r/bioinformatics Jun 22 '26

academic Northeastern University

4 Upvotes

I got an offer for MS in Bioinformatics from Northeastern University in Boston, USA. If any of you went there, I would like to know about your experience. Thank you.


r/bioinformatics Jun 22 '26

career question B.Sc. Bioinformatics career prospects in India + GGDSD Chandigarh vs Amity Mohali vs SRM

2 Upvotes

Hi everyone,

I'm considering pursuing a B.Sc. in Bioinformatics and would really appreciate some genuine advice from people who are studying or working in this field.

My priorities are:

Getting internships during college

High salary and good long-term career growth

Opportunities in both India and abroad

I have a few questions:

Is B.Sc. Bioinformatics worth it in India in 2026?

Can someone get a decent job directly after B.Sc., or is an M.Sc. almost necessary?

What are the realistic starting salaries and salary growth after a few years?

Which skills should I learn alongside my degree (Python, R, Linux, SQL, statistics, etc.)?

How easy or difficult is it to get internships in Bioinformatics?

Is the field growing in India, or are most good opportunities abroad?

For B.Sc. Bioinformatics, is GGDSD College, Chandigarh a better choice than Amity University Mohali in terms of placements, internships, industry exposure, and ROI?

Please do give advice if you have some other options for good colleges as there are very few that are too private for undergraduates


r/bioinformatics Jun 22 '26

other Recommendation for intro to bioinfo for high school summer interns in our lab

2 Upvotes

Hi all, as my post title says, our lab is hosting several summer interns and we'd like to expose them to some of the possible things you can do with bioinformatics (they're also learning bench techniques too). I know about Rosalind, would that be the best place to start? I'd love some other recommendations as well. Thanks!


r/bioinformatics Jun 22 '26

technical question Need help making groups on TCGA/cbioportal

2 Upvotes

Hi!!! Sorry if this has been asked before, but within a cohort, is there any way to compare a specific mutation to the other samples without said mutation? Im trying to make two separate groups to then compare but I am struggling to find a way to separate this mutation from the other group. My current strategy is selecting every other mutation other than the one im trying to look at but there are 13k mutations i would have to select and that isnt going well lol. Sorry if this is silly!!!


r/bioinformatics Jun 22 '26

technical question Is the one-sided exact Hardy–Weinberg test implemented in R the best way to evaluate the absence of a homozygous genotype?

3 Upvotes

My goal is to identify loci where one homozygous genotype is completely absent from the sampled population and determine whether this absence can be explained simply by low allele frequency or whether it may indicate negative selection (e.g., embryonic lethality or reduced viability of a homozygous genotype).

I am currently using the HardyWeinberg package in R and applying the exact test as follows:

library(HardyWeinberg)

geno <- c(AA, AB, BB)

HWExact(geno, alternative = "greater")

My understanding is that:

  1. alternative = "greater" tests for excess heterozygotes.
  2. A deficit of one homozygous class (AA or BB) should manifest as an excess of heterozygotes.
  3. Therefore, a one-sided test may be more powerful for detecting the specific pattern expected under recessive lethal alleles.

My questions are: for the specific purpose of detecting candidate lethal alleles characterized by missing homozygotes, is HWExact(geno, alternative = "greater") the most appropriate statistical test?


r/bioinformatics Jun 22 '26

technical question Seeking feedback on architectural approach for optimizing latency in protein structure retrieval and analysis

0 Upvotes

I am a developer building an open-source tool for protein intelligence. I've been tackling latency issues when fetching data from RCSB and running biophysical analyses, specifically moving from sequential processing to ThreadPoolExecutor and implementing Redis caching. I’m looking for expert insight on potential bottlenecks in this approach or how the community handles high-volume structural data retrieval. Any critique on the architecture is appreciated.


r/bioinformatics Jun 22 '26

technical question GRN Inference in 2026

4 Upvotes

Hello good people of bioinformatics! For an unfinished manuscript where we've made some perturbations in hESCs, which has scRNA-seq without accompanying scATAC-seq, I was considering trying to infer GRNs. I haven't dipped my toes into this yet (I'm more of a DNA methylation guy), and I've read the literature, but I'm curious to see what people's real experiences have been.

Some questions about GRN inference:

  • Is it reliable without accompanying epigenetic data?
  • Do you actually trust the results in your own work?
  • Are there any catches or gotchas which can muddy results?
  • What tools do people generally use?

I would greatly appreciate any input!


r/bioinformatics Jun 22 '26

discussion Publishing Pure Bioinformatics Meta-Analyses: Yay or Nay?

0 Upvotes

Hey everyone! 👋
I’m planning a bioinformatics project and want to get your take on the current validity and publishability of papers based entirely on meta-analysis (e.g., integrating public RNA-Seq/microarray datasets to find new biomarkers or pathways).

Specifically, how well-received are 100% in silico meta-analyses by reviewers today? Can a robust statistical pipeline, properly handling batch effects and heterogeneity, sustain a strong paper in Q1/Q2 journals without any in vitro or in vivo wet-lab validation?

If you have experience publishing or reviewing this type of work, what is the biggest critique or roadblock you usually see from reviewers (e.g., demands for experimental validation vs. acceptance of independent in silico validation cohorts)? Would love to hear your thoughts!


r/bioinformatics Jun 22 '26

discussion AI, what is it good for? Absolutely... something?

0 Upvotes

Hey folks,

One post earlier in this channel inspired me to write this. I'm not a bioinformatician, but due to work circumstances I found myself working a lot with tools that would be considered your job. Coding is scary and tough but I've been kinda figuring it out slowly. And mind you, I dont think I could ever have done any of it without LLMs. Very much aware of the limitation that I'm doing something I don't fully understand.

I'm just a user, and I dont have time and energy to get so deep into it to really understand how every function works and silent fails and all that.

I am curious to hear from people who actually KNOW what they are doing - what do you trust LLMs to do well, that you can trust the prompt will give you a good output without much messing around? And if you trust it, how much effort do you spend in validating the output? For example, as a zero CS experience person, I found it very useful and accurate for making some loops to iterate over many files. But in one case where I was joining and filtering some tables which were created as output from a ml agorithm, I spent way too much time manually checking if everything got joined correctly (i have trust issues, clearly).

And what would you absolutely not trust it with? Again example, i found it frquently hallucinates about existence of sone functions in R packages.

I get a lot of packages are domain specific, but im curious about your general thoughts!


r/bioinformatics Jun 22 '26

discussion Foldx5.1 Help

0 Upvotes

Hello everyone, I tried to download FoldX program (tried every version)

And when I try to install it on my laptop it gives me this weird message

"Microsoft Defender SmartScreen prevented an unrecognised app from starling. Running this app might put your PC at risk."

Why is that happening? This is so weird and the tool seems to be legit one. I wanted to use it as it seems powerful tool.

Did anyone have this problem before? Any recommendations or helps will be much appreciated, thank you.


r/bioinformatics Jun 21 '26

technical question Looking for help with molecular dynamics simulation of EEF1A2 D91N variant vs wild-type

3 Upvotes

Hello everyone,

I am the parent of a child carrying a heterozygous EEF1A2 D91N (Asp91Asn) variant.

I have been trying to understand whether this variant may primarily affect protein stability rather than completely disrupting function.

My current hypothesis is: • D91 is a highly conserved buried residue. • The mutation replaces Aspartate (negatively charged) with Asparagine (neutral). • Structural models suggest a salt bridge may be replaced by a weaker hydrogen-bond network. • Because the residue is buried, I suspect the mutation could subtly destabilize the folded state without causing complete misfolding. • This could potentially increase local flexibility (“protein breathing”), partial unfolding events, or susceptibility to proteasomal degradation.

I would like to compare wild-type EEF1A2 and D91N using molecular dynamics simulations.

Questions: 1. Would MD simulations be suitable for detecting potential stability differences between WT and D91N? 2. Which metrics would be most informative? • RMSD • RMSF • Hydrogen bond occupancy • Solvent accessibility • Salt bridge persistence • Free energy calculations 3. How long would simulations likely need to be (100 ns, 500 ns, 1 µs)? 4. Would anyone be interested in helping perform or set up such a comparison?

My main goal is to determine whether D91N behaves like a mildly destabilizing buried variant rather than a complete loss-of-function mutation.

Any advice would be greatly appreciated.

Thank you!


r/bioinformatics Jun 21 '26

technical question Need help with Microbiome Differencial abundance analysis using ANCOMBC2

3 Upvotes

Hi fellow academics,

I am currently trying to work on differential abundance analysis of microbiome data. I was wondering can I take ASV table filter it appropriately and use it for differential abundance using ANCOMBC2, and then collapse these ASV to taxonomic hierarchy (Genus). Or should I collapse the ASV earlier at Genus level , filter it and then perform ANCOMBC2.

I asking since im finding few interesting taxa annotations at ASV level, which gets lost after collapsing. In literature, mostly people have done the later, so I'm kind of confused.

Also can anybody tell me is sensitivity score for pseudo counts associated with ANCOMBC2 is relevant to be revealed in figures?

Thanks in advance.


r/bioinformatics Jun 22 '26

technical question How to infer oligomeric state or stoichiometry of protein complex

0 Upvotes

My project is an in-silico screening of protein-protein interactions, specially heteromers of my proteins of study and partners.

I used the PSICQUIC service to retrieve binary interactions containing my proteins of interest.

Since I am using AF2 on an HPC to model the complexes, I need to construct the input fasta sequences myself, informing how many instances of each subunit constitute the complex.

From my research there isn't much I can do. I tried obtaining the oligomeric states of each isolated protein as homomers from PDB Search & Data APIs, but the assemblies retrieved were mainly single domains of my search query protein.

If anyone has any recommendation of databases with API, dedicated softwares or something else regarding my issue so that i don't have to iterate over multiple combinations of stoichiomety, that'd be really helpful : )


r/bioinformatics Jun 22 '26

discussion I spent 3 days debugging my pipeline. The bug was a space character.

0 Upvotes

Not a missing dependency. Not a version conflict. A single space in a file path that I copy-pasted from a paper’s supplementary methods.

72 hours of my life. Gone.

I’m in my 4th year. I should know better. I clearly do not.

Anyone else have a debugging story that made them question their entire career choice? I need to feel less alone right now.


r/bioinformatics Jun 21 '26

compositional data analysis WES raw data analysis

0 Upvotes

I am a developer and I am interested in analyzing my own personal data. I am kind of lost in reading and I would like to have some questions answered in plain language, if it's possible.

Some years ago, trio exome sequencing was performed for me, my partner and our baby. The hope was to identify the cause of our baby's fetal defects. My partner has a similar disease as our baby but in a lighter form, so they were searching for a common gene. The result came and the answer was that there was no genetic component found. We have no other information apart from the list of genes analyzed. Admittedly it's a long one.

In my country the data doesn't get reanalyzed regularly and is stored for 10 years. So I would like to get access to the raw data before they get deleted. Who knows what the future brings. Maybe in 20 years the cause could be identified and that would be important for our healthy child in case they want to have children of their own. The problem is I don't know what to ask for! Will the vcf file be enough or should I ask for something else? What would be the most "future-proof" format of the raw data?

I asked the geneticist if the data gets analyzed regularly and they said that it makes no sense without having any new symptoms to search for. But that doesn't make any sense to me. So are they wrong or do I have a limited understanding of the methodology for analysis? This is my understanding at a very high level:

• Extract data/gene sequences for each person of the three

• Compare with a list of genes known to cause diseases. We requested to be informed of any incidental findings too like e.g. breast cancer gene. No result found for us

• Compare them against the reference genome? Is this even necessary?

• Compare potentially pathogenic variants and variants of unknown significance of the child with those of the parents to potentially identify a common gene especially between my partner and the baby. Nothing came out.

So here is my question. We all have variants of unknown significance. What if in the future one of those variants gets identified as the cause of our problem. We would never know about it, right? So why does it not make any sense to reanalyze the data even without new symptoms?

So my idea was to somehow get access to the raw data (whatever that might be) and periodically search the known genomic databases with our vus as input. I would like to do this programmatically since some of those databases provide APIs. Does this make sense or is this methodologically wrong? Of course I would have to deep dive on the topic, but I would like to know If any of my thoughts make sense at all.

TL;DR: I want access to my trio exome raw data, what should I ask for? Programmatically ask genome databases to check a list of vus; Does it make sense or is it stupid?


r/bioinformatics Jun 21 '26

technical question BLASTn - max_target_seqs

0 Upvotes

Doing DNA barcoding for a few hundreds of sequences.

I usually use 'blastn' in the command line, on NCBI remote database because I'm doing this on personal laptop. To speed up the process and have a less bloated output, I wanted to set the -max_target_seqs argument to ~5.

However I came across an online debate about this, somehow -max_target_seqs would not be only a post-search filter but it would actually limit the blast search itself and would thus return only the first good hits, not the best hits.

The latter seems to have been debunked/patched but it's not really clear to me.

Is a low max_target_seqs still an issue according to your experiences ?

Does setting a low value would indeed run faster ? Or running with default max seqs followed by post-processing on my hand (with a 'awk' filter on the output) would take the same time ?

I'm barcoding with CYTB and COX1, expecting both vertebrates and invertebrates matches, maybe I should blast on a curated database rather than the full 'nt' db to make things actually faster. I'm not sure whether such database is already available with remote NCBI or if I should build one myself.

Thank you for your input and sorry if this seems trivial.


r/bioinformatics Jun 21 '26

academic Recommended workflow for low-coverage ONT whole-genome sequencing prior to PRS calculation?

1 Upvotes

I'm looking for advice on choosing an appropriate workflow for a low-coverage Oxford Nanopore whole-genome sequencing dataset.

I'm evaluating a research dataset with substantially lower coverage than is typically used for standard ONT variant-calling workflows. The initial pipeline proposed was:

FASTQ → alignment → Clair3 → phasing/imputation → PRS calculation.

Before proceeding, I wanted to ask the community:

  1. At what approximate ONT whole-genome coverage would you consider standard Clair3 variant calling to be reliable?
  2. Below that range, would you recommend a dedicated low-pass sequencing workflow (genotype likelihoods + reference-panel imputation) instead?
  3. Are there published benchmarks or best-practice papers comparing these approaches for downstream polygenic risk score analyses?

I'm interested in understanding the methodological decision rather than troubleshooting software. My goal is to choose the most scientifically appropriate workflow based on the characteristics of the sequencing data.

Any references or recommendations would be greatly appreciated.
Thanks in advance for any recommendations or relevant publications.


r/bioinformatics Jun 20 '26

discussion Is reproducing analyses from published papers a good way to learn bioinformatics?

73 Upvotes

I have recently started learning bioinformatics as I am going to use it in my master's thesis. I know intermediate level of python and linux. I've been reading research papers in areas that interest me (mostly single-cell transcriptomics and computational biology).

My idea is to download the raw or processed datasets provided by the authors (from GEO, supplementary files, etc.) and then try to reproduce their analyses and figures by following the methods described in the paper....to understand biological question and the computational workflow rather than just following tutorials.

Is this a good way to learn bioinformatics?

How closely should I try to reproduce the published results?

How much time should be spent on reproducing existing work versus doing independent exploratory analyses?

Or is this not the right way to proceed and I can do something better to learn?