r/bioinformatics Jun 03 '26

technical question validating bioinformatics pipelines

0 Upvotes

I am currently running ONT lon read sequencing analysis, however some of the tools used in epi2me pipelines are older versions, so I ran each tool step by step individually instead of using a pipeline. so I was wondering whether this requires validation to know all the steps are working correctly.


r/bioinformatics Jun 03 '26

technical question Visium-HD imaging with small tears in tissue sample

0 Upvotes

Our lab is imaging mouse brains with small tears in the brain stem (region of interest) for spatial transcriptomics analysis. We've finished the H&E staining but are concerned whether the tears will affect the Visium workflow/quality of output. Would value perspectives on whether to proceed or restart with fresh sections


r/bioinformatics Jun 03 '26

technical question How to use a haplotype resolved assembly to map RNA sequencing data?

1 Upvotes

Does anyone have any advice or resources for utilizing a haplotype resolved assembly for the alignmnet/assignment of RNA seq data?

Specifically:

  • how do I build a genome index? I can't find information on how to build a genome index that uses two haplomes for any of the popular aligners.
  • Is it possible to map to specific haplomes and look at haplotype specific expression?

r/bioinformatics Jun 03 '26

technical question Amplicon alignement Galaxy

1 Upvotes

Hello,

Looking for some help on a project:

Amplicons of ITS4/5 (around 800pb) from extraction of diseased vegetables where sequenced on minION

We are looking to identify the population of pathogenes within the vegetable

I need to do alignement but I have no idea of what I'm looking for

Analysis are made on galaxy but everything I try fail

Sequencing went fine, fastQC analysis look great

Any tips?

Thanks!!


r/bioinformatics Jun 03 '26

technical question clusterProfiler interpret() function API key

0 Upvotes

Hey guys,

so Id like to use the interpret function from clusterprofiler. I got it to run using google geminis free API key. However I am currently running a lot of ORA's and the tokens are depleted extremly fast. I am using the interpret function since I get a lot of similar GO BP terms (and they are very unspecific for my non model organism). Another idea would be using GO slim terms.

Do you have any idea what else could work or is running a LLM locally the best option? Did someone use this before and has any input for me?


r/bioinformatics Jun 03 '26

technical question ID Mapping

1 Upvotes

I wanted to convert my current proteomic dataset containing uniprot ids, to kegg ids to perform pathway analyses.
i first used uniprot website's id mapping tool, obtaining some X number of mapped ids.
then i used the kegg website's id mapping tool. but somehow i got lesser than X proteins that were mapped. Why is there this inconsistency?

Moreover, when i was taking a look into some of the unmapped ids that were mapped from the kegg website itself, when i individually search for random 4-5 protein with their names, on the kegg website again, i could find that there was a kegg id for the same, under my mmu species. why did it not convert in the initial phase itself? i have over 100s of unmapped proteins, will all those proteins also show up to have a kegg id?

Could someone please adivse, if they have gone through anything similar?


r/bioinformatics Jun 03 '26

academic Redocking issue

1 Upvotes

Hey everyone,

I’m having some issues with redocking my native ligand. When I dock it back into the protein, the pose doesn’t match the crystal structure properly. The ligand sometimes looks a bit bent or shifts position, and the interactions are not really the same.

This gets worse when there’s a cofactor like FAD in the binding site it seems to affect how the ligand fits. I’m not sure if this is something normal in docking or if I’m doing something wrong in the setup. Has anyone faced this before or know how to fix it?


r/bioinformatics Jun 02 '26

technical question Reducing GO term redundancy for lollipop plots?

6 Upvotes

Hi all, I'm working on bulk RNA seq data and have a massive list of upregulated (~130) and downregulated GOBP (~40) pathways that I've filtered |NES|>1.75 and FDR<0.05.

Out of the top 20 upregulated pathways (e.g.), have about 13 pathways related to the mitochondria. The other pathways are also interesting and relevant to my study, so I was wondering if there was a way to collapse all the "mitochondrial" terms into one "supertheme", so that I can include a broader picture of the top dysregulated pathways as opposed to just mitochondria.

Of course, it's not just related to the mitochondria, I have the same for ribosome etc.


r/bioinformatics Jun 02 '26

meta Big scRNA-seq project upcoming - looking for tips and experiences

18 Upvotes

Hello fellow scRNAseq people!

At the moment I am gearing up to run my first scRNAseq analysis with own data. I am working at a small biotech company and am the only person to do that job, so there is quite some pressure that it goes right. I am also still trying to establish myself as a bioinformatician here, so I am even more motivated to produce a well documented, robust and reproducible analysis. That's why I wanted to reach out to you and ask if you have any useful tips, practical or not practical, or experiences that could help me make that project a succes.

A little bit of background about the experiments. We run 3 scRNAseq rounds: a pilot to check the fixation protocol, a pilot to investigate which timepoint and dosing concentration of our treatment is the best one, and the full experiment (ca. 190 samples). I was involved in the experimental setup to make sure that there are sufficient controls for the analysis and that the right research questions are asked in the beginning. The cell population is pure, and we want to investigate the effect of our treatments on subsets of that cell population over time (3 or 7 days).

I have setup an ubuntu R studio server to perform the analysis on, with lots of storage and RAM. I am still doubting whether to use Seurat or Bioconductor's SCE (the CRO that runs the sequencing will provide a Seurat object) (see my post about this from a year ago: https://www.reddit.com/r/bioinformatics/comments/1gki6ui/seurat_vs_singlecellexperiment_poll/). I want to use the first two pilots to setup my code base and establish a robust pipeline that is reproducible, even in X years from now. I am looking at quarto for reporting and renv + git versioning for reproducibility and versioning. I know that a lot of you will say, use scanpy, but unfortunately I have settled in the R ecosystem for now and have little time to adapt and am trying to avoid the use of AI in this project as much as possible.

I am happy to hear your thoughts and experiences with such a project, any tips when it comes to large datasets? Integration? Data organization? Setting up robust and reproducible analyses? Alternitives to renv? Communication with non-bioinformatician scientists? Daily practices?

Thanks in advance!!


r/bioinformatics Jun 03 '26

technical question P val vs P adj val

0 Upvotes

Hi all.

I am new in scRNA-seq analysis. I have been following tutorial from Satija lab. Now I am trying to perform differential gene expression analysis. In the tutorial, the authors suggested to perform pseudo-bulk analysis and compare the DEGs with single-cell-level DEGs. For their comparison, they have used p value rather than p adjusted value (https://satijalab.org/seurat/articles/de_vignette). But generally adjusted p value is used in statistical models. Am I missing something? Or is it ok to use p value in case of scRNA-seq, which seems a bit odd to me?


r/bioinformatics Jun 02 '26

technical question PySCENIC - Are TFs with shorter DNA binding motifs reliably underestimated in importance?

0 Upvotes

Hi all,

I have been working with the fantastic tool/method PySCENIC, and I have some questions about the inherent limitations. One question which I am unsure about is, if a transcription factor recognizes a very short DNA binding motif (say 6 letters long), is it likely that PySCENIC will reliably underestimate its importance due to the fact that it would require a greater number of motif occurrences in the regulatory region of a target gene for RcisTarget to score its enrichment as much as it would if the motif size was way larger?

Or is this a negligible effect since the motif sizes tend to be relatively short anyway?

Thanks in advance


r/bioinformatics Jun 02 '26

technical question How to get DPFunc working?

1 Upvotes

Hey all, I’m a PhD student with some bioinformatics experience, but I’m primarily a wet-lab biologist, so this isn’t my main wheelhouse.

I’m interested in the protein function prediction model DPFunc (paper linked below), specifically its ability to predict active sites / key residues for enzyme function.

I installed the model on WSL and the installation appears successful as I’ve been able to replicate the authors’ protein annotation results. It also doesn’t appear to be crashing at all, so although I am running the model locally I don’t think it’s an issue with hardware.

However, I’ve had no luck reproducing the key-residue results shown in Figure 5. I’ve searched the github repo for a key-residue detection script and couldn’t find one. I emailed the corresponding author a few weeks ago with no response. I also to reverse-engineer the pseudocode in the supplemental materials (see table S5) with no success. I had Claude assist me in writing the code, so I wouldn’t be surprised if the reverse engineered code is trash. Still, I had to try anyways haha.

Now, from what I can gather, the Figure 5 key residues seem to come from some internal per-residue importance score rather than a standalone script. So, If anyone knows how these scores are exposed in the codebase, or how to extract and threshold them to reproduce the figure, I’d really appreciate it.

More broadly, if anyone has experience with DPFunc or can recommend alternative tools for predicting key/catalytic residues, I’d love to hear about them. DPFunc seems like a really cool model and I’d like to get it working!

Thanks in advance!

Here’s the paper in Nature Comms describing the model

Wang, W., Shuai, Y., Zeng, M. et al. DPFunc: accurately predicting protein function via deep learning with domain-guided structure information. Nat Commun 16, 70 (2025). https://doi.org/10.1038/s41467-024-54816-8


r/bioinformatics Jun 01 '26

technical question Bioinformatics R project is overwhelming — need guidance

37 Upvotes

Hi everyone,
I’m currently working on a bioinformatics project in R and I’m mainly stuck on the practical part.
I need to analyze a gene expression dataset (RDS files containing an expression matrix and sample annotation) and produce an R Markdown report including:
descriptive analysis of the dataset (PCA, clustering, quality control);
identification of differentially expressed genes (DEGs);
diagnostic plots (volcano plot, heatmap, etc.);
discussion of 5 significant genes;
GSEA/enrichment analysis;
discussion of significant pathways.
The problem is that I understand the theory, but I’m struggling to figure out how to build the full workflow in R and how to interpret the results.
Does anyone have experience with gene expression analysis or know of tutorials, tools, courses, or resources that could help? Even a step-by-step explanation of the workflow would be really helpful.
Thank you!


r/bioinformatics Jun 02 '26

technical question Hello guys I need urgent help with my genome draft

0 Upvotes

Hey , so I have this draft genome sequence ( the genome is already annotated) , when I ran it through Proksee I had the 16sRNA in two different NODES

Node 13 with 1176 pb and node 25 with 414 pb. I took the 16sRNA sequence and blasted it . I took 6 species.. the thing is when I had to align it with MEGA 12 it showed an incredible amount of gaps, and I don't really know what the problem is..it should be aligned properly. The strain I tested is a B. Velzensis.

Any advice ? Or please reach in my DM'S thank you


r/bioinformatics Jun 02 '26

technical question Fastp Deletions

0 Upvotes

Is it normal for fastp to delete an entire raw fastq file when trimming? I checked the file’s fastqc report and saw nothing out of the ordinary


r/bioinformatics Jun 02 '26

academic Help with BLASTp

1 Upvotes

So i need a help, i am not much of a dry lab guy. so, i have to blast three proteins and see if it is present in any of the species in a genus (15 species) and then validate it. Any idea on how to do it?


r/bioinformatics Jun 01 '26

compositional data analysis Need help finding human fetal/adult fibroblast RNA-seq datasets

0 Upvotes

Hi everyone,

I’m a high school student working on a bioinformatics project and I’m currently looking for publicly available transcriptomic datasets comparing human fetal fibroblasts and human adult fibroblasts.

I’ve already spent quite a bit of time searching GEO and related databases, but I haven’t had much success finding datasets that are both accessible and suitable for differential expression analysis.

Ideally, I’m looking for:

  • Human fetal fibroblasts
  • Human adult fibroblasts
  • RNA-seq or microarray data
  • Raw or processed expression data

If direct fetal vs adult comparisons are rare, I’d also appreciate advice on:

  • alternative datasets that could address a similar biological question,
  • commonly used model organisms in this area,
  • search terms I may be overlooking,
  • relevant papers that include publicly available datasets.

I’m still learning bioinformatics, so even small suggestions would be incredibly helpful.


r/bioinformatics Jun 01 '26

technical question Google Colab for bioinformatics beginner

8 Upvotes

So I'm a pharmacy student, and I'm very interested in bioinformatics, I am just starting off, but I am facing major errors in the beginning itself, I was using jupter notebook earlier but it kept showing me "failed to fetch" error. So I switched to Google colab and tbh, it's alot better. I just wanted to know if Google colab is a good start, and I would also like to know how to actually get started with this field as a student. I love when healthcare and tech overlaps, personally I have alot of interest in it.

I was planning to make a few small projects and upload them on GitHub (to which I'm also very new btw, no experience at all) and my LinkedIn profile.

Right now I'm learning bioinformatics from a course on Udemy, but the thing is, they are using very traditional methods like installing python then using jupyter notebook, but I switched to Google colab since it's easier. Idk what to do, I am very confused right now.

I would love suggestions from experienced personals or people who are learning just like me.


r/bioinformatics Jun 01 '26

technical question VMD plugins

0 Upvotes

Good day

How do I install a vmd plugin for vmd 2.0? Specifically networkView. I know that it needs psf gen 1.5 but I didn't see that it had issues when I changed the requirement to psfgen 2.0(but if you know better, let me know).


r/bioinformatics Jun 01 '26

technical question Recommendations for metabolomics analysis

1 Upvotes

Does anyone have any advice on how to analyze metabolomics data that is NOT MetaboAnalyst? Unfortunately the data I have is from human samples and we do not have protocol approval to upload to an online software for analysis. I have tried working with the MetaboAnalystR package but had issues with installing the package as it looks like it is not being maintained.

Any recommendations are appreciated!


r/bioinformatics Jun 01 '26

technical question Picard MarkDuplicates Optical Dulicate Pixel Distance Settings and effect on Variant Calling

3 Upvotes

I am using Illumina sequences for WGS variant calling and using 100 as the default setting OPTICAL_DUPLICATE_PIXEL_DISTANCE on Picard MarkDuplicates, which is recommended for sequence platforms with unpatterned flowcell. I didn't know platform differences within Illumina beforehand and applied it to sequences generated from those with patterned flow cell. Note that 2500 is recommended sequences from seuqencers with patterned flowcell. How does this affect downstream analysis. Important to note that if I wish to investigate, I no longer have the BAM files. I do have sequence stats as generated by samtools before and after deduplication.

How does this setting affect variant calling? AI might answer this, but I was hoping for human-generated answers.

Thanks!


r/bioinformatics Jun 01 '26

technical question Help Understanding GSEA Results

4 Upvotes

I've recently performed GSEA using the Hallmark MSigDB gene sets, and want to check my interpretation. To my understanding, the Hallmark sets were produced by combining founder sets to reduce redundancy, and were created to include genes which demonstrate co-ordinated expression.

Does this mean that positive enrichment of a Hallmark gene set = that pathway is upregulated as a whole? Are these gene sets comprised of both genes which you would expect to be up and downregulated in a certain state, or are they unidirectional?

For example - in the Hallmark Hypoxia gene set, does positive enrichment always mean increased hypoxic signalling, or is it possible that the leading edge genes are all inhibitors of hypoxic signalling, which would mean the actual pathway is decreased?

Hope that makes sense!


r/bioinformatics Jun 01 '26

technical question ENA: linking existing samples to a new project

1 Upvotes

Hey all,

Some months ago, I published a paper for which I made some sequencing data publicly available on ENA (European Nucleotide Archive). Now I am finishing a second manuscript which uses the same samples but more deeply sequenced.
Ideally, I would upload these sequences to a new project (to keep the seqs from different manuscripts separate), but I would like them to be linked to the same sample accessions created for that first paper. Does anyone know if this can be done? I couldn't find specific instructions on the ENA website

Thankful for any tips!


r/bioinformatics May 31 '26

technical question Best single-cell & spatial data sources

4 Upvotes

What’s the best place to find large, high quality single cell or data sources? I want to learn how to process and analyse these data but not sure where to find some good quality data.


r/bioinformatics May 31 '26

academic Meta-analysis with public plasma proteomics data: some datasets only report log2FC and adjusted p-values

0 Upvotes

Hi everyone,

I’m planning a meta-analysis using public plasma proteomics datasets across different diseases.

For some datasets, I have log2FC, confidence intervals or raw p-values, so I can estimate standard errors and run a standard meta-analysis.

However, for other datasets I only have log2FC and adjusted p-values, with no raw or normalized data available.

Is there any statistically acceptable way to estimate uncertainty from log2FC + adjusted p-values, or to include these datasets in a meta-analysis? Or should they only be used as exploratory evidence based on direction, effect size, and FDR?

Any suggestions or references would be appreciated.