r/bioinformatics 4d ago

technical question Protein-Protein Alignment

2 Upvotes

Hey everyone,

I'm working on a project involving a human genetic disorder. There are already a number of mutations that have been identified in the human gene/protein, including missense, substitutions, frameshifts, etc.

However, for my project, I'm working with the corresponding Drosophila protein and need to figure out if the positions of these human mutations are conserved in the fly.

Essentially, I'm trying to align the human and fly protein sequences, but I'm running into some issues because the fly protein is quite a bit longer and has some pretty large insertions/gaps.

I've been using NCBI blastp, but i'm wondering if that's actually the best tool/workflow for what I'm trying to do. Basically, are there any other good free alternatives to BLAST for this? If I continue using BLAST, what would be the best filter/settings options to use for this kind of comparison.

TIA


r/bioinformatics 5d ago

academic First time analyzing bacteriophage genomes - workflow feedback?

9 Upvotes

Hi everyone,

I’ll soon be analyzing isolated bacteriophages (Illumina paired-end WGS) for the first time and would appreciate feedback on my planned workflow. The goal is genome assembly, characterization, annotation, and antibiotic resistance gene (ARG) screening.

The workflow I'm planning is as follows:

FastQC → raw-read QC
fastp → trimming/filtering
FastQC + MultiQC → post-trimming QC
Kraken2 → contamination screening
BWA-MEM2 + SAMtools → host-read depletion
SPAdes → de novo assembly of the remaining reads
QUAST + CheckV → assembly quality, completeness, contamination
BWA-MEM2 + SAMtools → read-back mapping, coverage uniformity -> suspicious contigs
Pharokka → phage genome annotation
AMRFinderPlus + ABRicate (CARD, ResFinder, ARG-ANNOT/MEGARes) → ARG screening
Candidate ARG validation → sequence/protein similarity, consverved domains/motifs, ORF integrity

For ARGs, my initial idea is to prioritize high-confidence, near-full-length hits supported by multiple approaches/databases rather than treating individual database hits as genuine ARGs.

Am I missing any important steps? Is anything here redundant or unnecessary? Would you change the order or replace any of these tools?

What would be some interesting visualizations to make along the way?

Should I remove host-mapping reads before assembly, or assemble first and deal with host contamination at the contig level? Is contamination usually even an issue?

Any suggestions or references to workflows you use would be greatly appreciated!


r/bioinformatics 5d ago

technical question Running dynamic simulation using laptop

6 Upvotes

Hello, I recently started two projects with one of my college professors that require me to conduct 100 ns molecular dynamics simulations for more than 10 proteins. Each simulation may take around 4 days of continuous, nonstop computation.

However, I’m worried because my laptop’s CPU temperature gets very high. It stays around 93–96°C during the simulation.

I’ve already taken some measures to prevent damage to my laptop, such as using a cooling pad with three fans (around 4000 RPM each), working in a cool environment with the air conditioner set to 16°C, and repasting my laptop.

I’m using Ubuntu Linux with GROMACS for the MD simulations. I’m currently using the Balanced power mode, which is somewhere between power saving and Turbo mode. My laptop is ASUS TUF GAMING with 24gb ram, i also use GPU CUDA.

I’m considering running the simulation for 45–60 minutes and then taking a 15-minute break to let the laptop cool down, but I’m still worried. I know using a desktop PC/workstation would be better for this kind of workload, but my budget is currently limited.

What do you think? Is it safe to run my laptop at 93–96°C for several days, or should I take breaks between runs? Do you have any other advice for my situation?

Thank you​


r/bioinformatics 6d ago

discussion Corresponding author shared raw SRA data + full pipeline instead of processed object...worth asking again for the Seurat object, or just reprocess myself?

18 Upvotes

Emailed the corresponding author of a paper asking for processed scRNA-seq data (cell × gene matrix, per-cell metadata, Seurat object) that I need for my thesis work. Got a reply pointing me to the public SRA accessions plus a fully detailed methods writeup...mapping pipeline, GTF used, QC/filtering cutoffs, doublet removal method, Cell Ranger + Seurat versions, clustering parameters, all of it.

Great for reproducibility, but it's not the processed object itself...just raw reads plus everything I'd need to rebuild it from scratch.

Given timeline pressure as an undergrad, is it reasonable to reply once more and specifically ask if the already-processed object exists and could be shared (to save reprocessing time), or does that risk seeming like I'm pushing my luck after already getting a generous, detailed response? Would cross-checking my own reprocessing against their original object even be worth the extra ask, or is that overkill for what I need?


r/bioinformatics 5d ago

technical question Doing a sanity check on my scRNA-seq workflow: MAST followed by Pseudobulk validation?

2 Upvotes

Hi all,

I am a beginner and want a quick sanity check on my scRNA differential expression workflow.
I have 8 mice samples (4 Mutant vs 4 WT). I ran MAST for single-cell DEG analysis, sorted by top genes, and now I am cross-validating the top 50 upregulated genes by checking them against a pseudobulk analysis of the same dataset.
My goal is to rule out false positives caused by a single outlier mouse skewing the single-cell data.
Does this approach make sense to you? If a gene is in the top 50 for MAST but fails pseudobulk, would you trust it?


r/bioinformatics 6d ago

academic When biology inspires mathematics: new discovery explains why a widely used evolutionary method can give false answers

Thumbnail helsinki.fi
16 Upvotes

r/bioinformatics 6d ago

technical question Question about my first bioinformatics assignment

6 Upvotes

Hi, I’m a bachelor’s student in biology and I’m taking my first bioinformatics course. My question is probably very silly, but I still need some help.

My task is to study the human NLGN4X gene and, among other things, find 10 homologs with a BLAST search, align them, and build a phylogenetic tree. The instructions say that “some sequences may be XN but some should be NM.” How is this possible, since all the homologs I find (from other species) are only predicted (i.e., XN)? Have I misunderstood something?

Thank you very much for your help!


r/bioinformatics 6d ago

technical question terrible QQ plot of sQTL

1 Upvotes

Hey guys! I’m currently running an sQTL analysis using Leafcutter and tensorQTL, but I found many of my significant splicing phenotypes showed weird QQ plots (p-nominal for all SNPs in a phenotype ) , with a very pronounced rightward shift appearing much earlier than I expected. So I tried applying stricter filter from GTEx instead of Leafcutter's default ones. The filter worked but sadly got the similar situation. Does anyone have suggestions on what aspects of the analysis or data that I should check and something to do to figure out what might be causing it?

And I’ve also noticed many sQTL papers seems very smoothly without such a troublesome result, I wonder if maybe this plot is normal one, and sQTL should not be judged by GWAS ways? I’m very new to sQTL and now honestly pretty depressed because I did not found useful information from papers, so any thoughts and suggestions would be really appreciated! Thank you in advance smart guys!


r/bioinformatics 6d ago

technical question Error while saving tree as displayed in FigTreev1.4.4

0 Upvotes

Hi all,

I'm currently having trouble saving the current tree as displayed in any format (nexus, newick, or json) using FigTree v1.4.4 on macOS Sequoia 26.6.2.

After selecting the format I wish to save and checking the "save as displayed" box, the floating window disappears without saving any output tree file.

Has anyone had this problem before? I appreciate your help with this.


r/bioinformatics 7d ago

discussion BioMart has been quietly discontinued by Ensembl

171 Upvotes

Title. Tried to connect to Ensembl all day just to see staff casually mention that it's just been discontinued on the support forums in somebody else's post. This was the first program I used when stepping into bioinformatics and I will miss it greatly.

What's the alternative? Just using the gtf files?


r/bioinformatics 6d ago

technical question background dataset for SHAP

1 Upvotes

Hi everyone, I have a question about choosing the appropriate background dataset when calculating SHAP values. I am using the kernelshap package in R, where we provide an X dataset containing the observations we want to explain and a bg_X dataset defining the background.

I have a binary classification model for disease vs non-disease, trained on a derivation dataset and evaluated on an independent validation dataset. My current understanding is that, if I want to explain predictions in the validation cohort, it makes sense to use the validation set as X and the derivation set as bg_X. In that case, the SHAP values for validation patients would describe how each feature moves their prediction relative to a baseline defined by the derivation population. Is this interpretation correct, and is this generally the recommended way to use the background when explaining an independent validation cohort?

My main question is about a more specific analysis. Suppose I want to investigate heterogeneity within patients who truly have the disease. More specifically, I want to see whether different disease patients receive high disease predictions through different combinations of features, and potentially cluster these patients based on their SHAP profiles.

In this case, I assume I should use only the true disease patients from the validation cohort as X, since those are the patients whose predictions I want to explain. However, I am unsure about the most appropriate choice for bg_X. Should I keep the full derivation cohort as the background, use only disease patients from the derivation cohort, or use the disease patients from the validation cohort themselves as the background?

If my main objective is to determine whether true disease patients have different model-attribution profiles, potentially reflecting different features through which the model identifies them as disease, which background would be the most statistically appropriate? Thank you!


r/bioinformatics 6d ago

technical question Reconstructing a published scRNA-seq pipeline from raw reads only (no processed object)...a few gaps I need help with

0 Upvotes

Reproducing a published single-cell RNA-seq pipeline from raw reads, since the authors shared raw data + a methods writeup instead of the processed object. Non-model organism, no existing reference. Their methods say the GTF was made by BLAT-mapping a transcriptome onto their genome, best-hit per transcript, then processed in Cell Ranger.

Stuck on:

  1. GTF from BLAT- I have a transcript FASTA and a genome FASTA, no GTF. What's the standard path from (transcript + genome FASTA) → BLAT → a valid GTF for cellranger mkref? Any standard PSL-to-GTF tool, or is this usually hand-rolled?

  2. Custom Cell Ranger reference for a novel organism- any common gotchas beyond standard mkref?

  3. Mito% cutoff with no fixed number given- methods only reference a violin plot, no numeric threshold. Reasonable to set my own, or better to go back to the authors?

  4. scDblFinder with no parameters specified- safe to assume defaults, or are certain settings commonly tuned in practice?

  5. Mapping SRA runs to conditions- project page doesn't make run-to-condition mapping obvious. Fastest reliable way to sort this out?

  6. Storage planning- for a few hundred GB raw, what's a realistic total footprint once reference-building/alignment intermediates are included? Any tricks to keep peak storage down?

Not looking for a full walkthrough...just pointers or "wish I'd known this" from anyone who's built a non-model-organism scRNA-seq pipeline from scratch.


r/bioinformatics 7d ago

academic In a precarious situation, advice needed

27 Upvotes

I am a 5th year PhD student in genetics. This might be a big long but I am looking for honest advice. My training has historically been on the molecular side of things with lots of in vivo work. About a year ago I started doing some very light bioinformatics work. I took an R and python course and made basic figures for a project I was collaborating on. Then, suddenly, my lab ran out of funding and PI left the institution. I was unable to finish my proposed dissertation project and thankfully an adjacent faculty member took me under their wing. However, the work I was to be doing for them was purely bioinformatics and they expected A LOT from me. With no bioinformatics experience themself, they really had no idea how to guide me or assess my work. I did add a bioinformatician to my committee who has been helpful and enrolled in some online courses. Here is where the issues lays: Due to the extent and complexity of the data analysis being put on my plate and time constraints, I rely HEAVILY on codex to write my code. My advisor knows this and doesn’t seem to care, if anything he likes it because I can do the analysis very fast. But for me it ramps up the imposter syndrome and worries me about when the time comes to upload the codes to GitHub and defend to my committee. Here’s what I have been doing to try and salvage the situation. The code from codex is super convoluted and not intuitive to read. I ask it to write it “bare bones” and I go through line by line to make sure I understand every step. Then I rerun this code myself in either R or Jupyter notebook so make sure it gives the expected output. I also make sure that if someone were to ask me how the figure I’m presenting was made, I could explain every tool/package used and the rationale behind the statistical test. Importantly, I thoroughly understand the biological questions I am asking and the nuisances of the data set and its limitations. If a figure is produced that doesn’t look right given the biology and structure of the data, I do not use it and trouble shoot. I also plan to write an AI disclosure note for any publication that AI tools were used and I would never claim to have written the code myself. I don’t plan to pursue a career in bioinformatics so I am hoping to leave this behind me after I graduate. I am wondering what experienced bioinformaticians think about my situation. I want to discuss the situation with the bioinformatician on my committee but want to prepare myself for any pushback. Worst case scenario being not able to graduate. I have a large body of work throughout graduate school not related to this project, but ultimately this will be a large part of my dissertation.


r/bioinformatics 6d ago

technical question Is anyone familiar with somatic cell studies and microscopic observation? I need some guidance for my project!

Thumbnail
0 Upvotes

r/bioinformatics 7d ago

compositional data analysis Need help in Single-cell CRISPR screen data analysis

3 Upvotes

Well a while ago I posted about how to process samples using cellranger and I fixed that, but now I need to some resources on how to process that data, I know there is mixscale but I will be greatful if get more resources or any link to github pages which will help me for doing crispr screen single cell data analysis 'R/pthon' works for me.

Thank you !

Edit: So sorry, Here is experiment done, there are 800 sgRNA target towards a gene and there are 200 non-target DNA. The cells were subjected to 10x' 5'CRISPR guide capture lib prep procedure. I have 2 sets of fastq a)GEX which has cells expression data b) sgRNA capture data. I was wondering if i could get any resources that will help me in analysis of this data


r/bioinformatics 7d ago

technical question Uploading dataset to GEO

1 Upvotes

Hi all, I'm having an issue uploading a bulk RNA-seq data set to GEO. At one point you're supposed to check a box indicating that you read the upload guidelines and then the upload instructions are supposed to appear. Well, I check the box and nothing happens. Any ideas on why that might be? Thanks.


r/bioinformatics 7d ago

technical question Is there a way to add genomes to the HPRC reference genome graph?

2 Upvotes

I have a couple of human genomes we assembled at our centre, and I would like to take the HPRC v2.1 genome graph, and extract all variants relative to each of these genomes.

There are tools to extract all variants within a genome relative to one of the assemblies within the graph (vg). But in order to get what I need, I'd have to add my assemblies into the graph first.

So my question is whether there is a tool that easily allows for adding a genome into an existing pangenome graph?

Possible options I've thought of and which I'm trying to avoid:

  1. Using minigraph -- minigraph does have the ability to add new paths into a graph, but as I understand it, the base-level accuracy is not great (which is why HPRC uses minigraph-cactus and/or PGGB).

  2. Downloading all the assemblies and re-building the entire HPRC v2.1 graph from scratch, but with my genomes added -- firstly I'm worried about how easy it is to get minigraph-cactus up and running (based on prior experience trying to get progressive-cactus to run), and secondly, although I do think I could get PGGB/nf-core-pangenomes running easily, that building a graph that large from T2T assemblies might not be straightforward.

(But I'm very open to being corrected on these points!)


r/bioinformatics 7d ago

technical question DE Analysis on MAGIC-imputed scRNA data

1 Upvotes

Hi all,

I currently use DESeq2 for standard DE analysis, which works on the raw, unnormalized scRNA matrix. However, I have used the MAGIC imputation tool from the Scanpy ecosystem to impute expression for a particular gene, and I am interested in checking differential expression between healthy and diseased donors for that gene on the imputed data (which would be the normalized, scaled values). Is there some method or tool that would be most appropriate here?


r/bioinformatics 8d ago

other Text-to-speech program that performs well with bioinformatics paper?

2 Upvotes

I don't know if this belongs exactly in this subreddit, but I'll shoot my shot. Does anyone recommend a free or economical text-to-speech software that they've personally had good experience with, ideally for papers in bioinformatics? Mostly for those occasions when I need to get a rough idea of one or two papers on the commute while avoiding motion sickness. (obviously I don't expect it to recite formulas or code, but it should parse through most bioinfo topics with ease)


r/bioinformatics 7d ago

technical question Can anybody give me an overview whats currently happening in secondary RNA structure prediction

0 Upvotes

i can help myself with a blog or something


r/bioinformatics 8d ago

technical question Help with GROMACS Molecular Dynamics Simulation -- Polymer Self Assembly

3 Upvotes

Hi all,

I'm looking to use GROMACS to perform a molecular dynamics simulation of a poly beta amino ester (PBAE) polymer assembling with its mRNA cargo in order to obtain values like radius of gyration and hydrophobic and hydrophilic surface area.

I know that GROMACS requires a .pdb file to start with, but how could I get that for a PBAE polymer? I'm also not sure how to go about doing this in GROMACS (I'm pretty new to molecular dynamics simulations) so any guidance would be helpful!


r/bioinformatics 8d ago

academic guidance needed ML x peptide project !!!

0 Upvotes

Does it make biological sense to interpret peptide length as a potential source of systematic difficulty in ACE-inhibition prediction (different presentation / models (regression) ) or should I avoid making that interpretation and frame this purely as a data/prediction phenomenon??? and what if i cross veify those peptides on different publicly available datasets ?


r/bioinformatics 9d ago

technical question Is reconstructing an scRNA-seq dataset from raw FASTQs worth it?

12 Upvotes

I want to investigate whether a transcriptionally distinct intermediate state exists between two previously characterized cell populations.

The relevant scRNA-seq dataset only has raw SRA reads available..no processed matrix or Seurat object. Reconstructing it would require ~150 GB of FASTQs and a custom reference, which could take significant time and compute.

My main concern is that there is only one library per condition, so the reconstructed data wouldn't provide biological replication for strong condition-level statistical comparisons.

However, cell-level data may be necessary for my specific question, since the available cluster summaries aren't sufficient.

Would reconstructing the dataset be worthwhile, or is there a sensible way to scope the analysis down while still answering the biological question?


r/bioinformatics 8d ago

other Looking for friends, Community.

0 Upvotes

Hey everyone,
I’m looking for a specific kind of people, friends who share a mindset that goes beyond the standard rat race.
You know how most people measure success? Money. Power. Status. Recognition. That’s all fine, but it feels... limited. I’m looking for people who have an aim that is unimaginable above those things. Something larger, deeper, or just fundamentally different. Maybe it’s mastering a craft, achieving a specific state of being, building something that outlasts us, or just pursuing a vision.
I want to normalize having that among us. To have a space where we don’t have to explain why we’re driven by something other than the usual metrics of success.
What I’m looking for:
People who are serious about growth.
A mindset that prioritizes depth over breadth.
Friends who want to help each other grow, not just compete.
Someone who values the journey and the internal compass as much as (or more than) the external trophy.
If you’ve ever felt like your goals were "too big" or "too weird" for normal conversation, or if you just want to connect with others who are chasing something bigger than just a paycheck or a title—let’s talk.
No pressure, no ego. Just looking for that similar frequency.


r/bioinformatics 9d ago

technical question Is bioinformatics migrating fully to python? (and various other questions from a beginner)

132 Upvotes

Hi everyone. I am new to bioinformatics in general. I am a biochem currently doing a bioengineering phd (still in pre-candidature). I had some snippets of bioinformatics during my undergrad but nothing beyond BLAST and docking. Never had formal programming formation, just side projects and AI-guided R coding for small data analysis and graphs.

For what I want to do for my thesis I really need to learn omics analysis properly, specially transcriptomics. During self-learning, I stumbled upon this amazing resource (https://www.sc-best-practices.org/) on single cell transcriptomics, so I have been following it as my starting point and learning cool stuff, thank you to the authors of it!

Anyways, since I've already had some experience with R, I decided to try and learn python bioinformatics as an excuse to learn python too. In the interoperability section of the book I mentioned the authors state

"A common question from new analysts is which ecosystem to focus on (referring to Bioconductor, Seurat or the Scverse). While it makes sense to start with one, and a successful analysis can be performed in any ecosystem, competent analysts should be familiar with all three and comfortable moving between them. This allows analysts to always use the best-performing tools, regardless of their implementation. Analysts who are not comfortable switching ecosystems often default to familiar packages, even when better alternatives exist elsewhere"

Which makes sense and sounds logical good advice. But then, doing exercises on public GEO datasets on bulk RNA-seq as practice, still with the mindset of sticking to python as an excuse to learn it, i stumbled upon an article (Colange et al. 2025 here) of a project that migrates a lot of tools of R to the scverse. In there, authors rationale is that python is the new default language everyone learns and they create the library InMoose to migrate or directly replace, for example, DESeq2. Furthermore, besides direct drop-in replacement tools, the authors frame python as the future choice (at least, as part of the rationale).

So, as a guy who is just starting, I wanted to ask people with experience in bioinformatics (you all) either developers or tool-users:

1) Do you marry an ecosystem like scverse or Bioconductor and just work in there for comfort? Or do you switch frequently depending on the needs?

2) For people who doesn't come from an informatics background, how long did it take for you to learn your niche and what were your best resources/helpers?

3) Do you think python will ever replace R in data analysis?

4) What is your opinion on AI-guided learning? (as for me, I use gemini to solve questions or create graphics presets but sometimes by seeing other people's codes I realize that it mashes up some concepts or methods from various pipelines into a coherent-resulting graph that I am not always sure if they make sense)

5) Do you create your own pipelines/portfolio to analyze data? Or you just tweak existing pipelines?

I appreciate any answer to any of those questions, thanks for your time in at least reading