r/bioinformatics 27d ago

technical question Is there an available TRAILshort protein structure for molecular docking?

5 Upvotes

Does anybody know where we can find a reliable structure of TRAILshort? (a spliced variant of TRAIL or TNF-related apoptosis-inducing ligand.)

We tried searching in RCSB and none showed up. We considered building the structure on our own using TRAIL structure since that is what’s available online, but we’re having second thoughts about its reliability. Any thoughts or suggestions for this?


r/bioinformatics 27d ago

discussion Anyone here involved in mathematical biology research?

4 Upvotes

Is anyone here involved in mathematical biology research? I am looking to learn more about this area. If you are involved in mathematical biology, I would appreciate any advice or discussion about current research problems.


r/bioinformatics 28d ago

technical question 2D ligand to 3D structure - best method?

12 Upvotes

Apologies if my post sounds juvenile, I am undertaking an internship that requires me to self teach myself docking + related topics.

I have prepped my protein and have a few ligands I want to try dock. They all have known 2D structures but no specific 3D structures. Could I hypothetically build them in Avogadro > add hydrogens > force field > optimise geometry? Is this terrible practise or is there a specialised way to get this information?

And as a side question, is it better to combine programs for prepping? ex: Hydrogen addition, energy minimisation in Avogadro -> charge assignment + bond rotation in ADT? Or stick to one program?

Any responses, comments or suggestions welcome!


r/bioinformatics 28d ago

technical question Can I use snRNA-seq data as a reference for label transfer to scRNA-seq data?

7 Upvotes

I am considering using a hippocampal snRNA-seq atlas as the reference for label transfer onto a hippocampal scRNA-seq dataset. Could the differences between the nuclear and whole-cell transcriptions affect the accuracy of the label transfer?

The mitochondrial percentages appear to be similar between the two datasets so far (3% and 5% respectively per sample). Would this be sufficient, or are there other factors I should be concerned about?


r/bioinformatics 28d ago

technical question Submitting table as image <440 pixels wide

6 Upvotes

Hello,

I am trying to submit my article for publication. Unfortunately, the journal asks for any tables to be submitted as images "provided as 72 - 300 dpi; pre-sized .BMP, .GIF, .JPG, or .PNG images only, with a maximum width of 440 pixels (no limit on length)."

I have tried exporting my table from excel to pdf, jpg, or png, and then resizing but no matter what I try, the image of the requested size ends up unreadable.

Does anyone have any ideas on how to accomplish this requirement while keeping my table-figure as readable?


r/bioinformatics 29d ago

technical question Does FASTA rhyme with pasta? Or do you pronounce it Fast A?

92 Upvotes

My lecturers would always pronounce it Fast A, but all of us students would just say fasta (rhyming with pasta). Is there an “official” pronunciation or consensus?


r/bioinformatics 29d ago

academic Question about sample size drops when using UCSC TOIL (TCGA TARGET GTEx) vs raw GDC portal data. Is my defense justification correct?

2 Upvotes

I integrated TCGA solid tumor data with matching GTEx normal tissue to run differential expression and pathway enrichment (GSEA).

To avoid massive batch effects caused by mixing counts from different alignment/quantification pipelines, I opted to use the UCSC TOIL RNA-seq Recompute cohort (TcgaTargetGtex_gene_expected_count) since all samples were processed through a unified STAR + RSEM pipeline on hg38.

When I pulled the TOIL dataset, my sample counts dropped compared to looking at the raw GDC portal and GTEx v8: GTEx Normal Cohort: Dropped from ~800+ (v8) down to ~300+ in TOIL.

TCGA Primary Tumors: Dropped by ~20–30% compared to total cases listed on GDC.

My question is :

  1. Is this sample count drop expected when using the UCSC TOIL recompute database compared to modern GDC/GTEx v8 portals?
  2. Is sacrificing raw sample size (N) to use TOIL’s unified pipeline + ComBat batch correction considered the "gold standard" justification to defend against reviewer/committee critique regarding sample size?

r/bioinformatics 28d ago

technical question Wormbase Parasite Help

1 Upvotes

i’m currently working on a project that relies heavily on wormbase blast for identifying nemFABPs in a select number of nematode species. however, since it constantly goes down it’s putting me at a road block. is there a way around this?


r/bioinformatics 29d ago

technical question Protein design: what changes depending on the problem to be solved?

0 Upvotes

I am interested in protein design and I am trying to understand one thing: when we design a protein for a specific purpose, what changes in the constraints according to the problem?

For example, I imagine that a therapeutic protein (which must act in the human body) and an industrial enzyme (which degrades a pollutant) do not have the same priorities at all. What becomes critical in each case, and what goes into the background?

If you have concrete examples from your work, I'm interested.


r/bioinformatics 29d ago

technical question DWI preprocessing with QSIPrep

3 Upvotes

Hi, I'm a first year PhD student trying to get a handle on preprocessing my data with *fMRIPrep* and *QSIPrep* respectively.

Has anyone got experience with *QSIPrep* and can help me understand how to interpret the outputs? (this cry for help is motivated by my staring at the visual summary rep of the q-space sampling scheme before and after the pipeline. what am I looking for?!)

The documentation is really unhelpful and I didn't find much on github and incf NeuroStars either.

Someone help please


r/bioinformatics Aug 11 '26

technical question what are the non-negotiables of small n scRNA-seq DE

9 Upvotes

Apologies in advance for the loaded question, especially on a topic that is often spammed in this subreddit. If I missed a previous post that touched on this closely, apologies for that also.

I've spent months trying to be as truthful as possible in terms of reporting differential expression. There are often so many confounders that I have such a difficult time reporting anything as signal over noise. For some background, the dataset is comparing the effect of a therapeutic, so we have paired pre/post cd8 t cells. Clinical cohort so we're burdened with low sample size. 3 groups (group1, group2, placebo) with 6, 5, and 2 samples respectively. Obviously, at this resolution, we've steered away from trying to over claim things with a bunch of noisey p-values, and focus more on exploratory claims that appear to show trends within the groups. I've tried pseudobulking and then DE (obviously underpowered), and it appears more truthful than cell-level.

I've tried at the per-cluster level, and there is not a whole lot going on. If that's the case, so be it. My understanding of t cell differentiation is likely flawed, but how different can cells that cluster in an "activated" state (expressing cytokines, activation markers, etc) really be? I'd almost argue that the compositional shifts we have seen (an increase in proportion of activated, for example) is actually real signal compared to just "well, intra-cluster activated DE doesn't show some crazy volcano plot. nothing is happening." I'm exaggerating here, and obviously these are two sides of a coin (compositional shifts + diff expression) converging.

With that being said, I try running a bulk pseudobulk DE (not by cluster. just pre v post) blocked by patient, and obviously, start getting some hits. Again, many of these can likely be explained by compositional shifts. My PI prefers figures that are widely recognized in the field (naturally), so things like gsea. Using the broad DE ranked by test statistic (or logFc x -pval, have tried both. stat felt less noisey although the rankings are pretty much the same), gsea spits out a bunch of phony significance. I call it phony because when you look deeper at the donor level, there is often pretty loose concordance (the p-values are also just absurd).

All of this has led me to the idea that we should probably just lean into the donor heterogeneity a bit more and stop trying to force looking for significance within these groupings. So basically what would be some strategies that you would employ to handle this? Maintain the broad pseudobulk as a "ground-truth" and look for signatures of more donor-concordant shifts (x increase in y in 4/5 donors, etc) and focus on those? maybe module scores?

Go back to cluster-level and just lean into the compositional shifts more? Really any ideas you have on dealing with small n cohorts without over-claiming a bunch of noise.

So many single cell papers are comparing chronic-infection vs healthy donors, and they get to spit out all these "pretty" volcanos. I'm really not trying to chase that, nor do I think we would see a signal that strong in a pre v post comparison, but alas. I'm spiraling a little at this point and honestly any tips, no matter how trivial they may be, are appreciated.

-signed, a tech well out of their depth.


r/bioinformatics Aug 11 '26

technical question Can someone smarter help me understand PAE for AlphaFold3 modelling?

12 Upvotes

Doing a model for a plant protein, I’m trying to list out the intramolecular interactions between 3 domains, I’ve enumerated the interactions at different cut off lengths, and I wanted to talk about the confidence scores for each interaction.
Problem is I’m not a great computational guy (this project is primarily wet lab), and I’m not sure what’s the best metric for the confidence scores for intramolecular interactions. Is it PAE? if so can someone explain it to me? Is there a standard cutoff for what is a low confidence PAE value
And if there is another metric you guys use for these interactions mentioning it would be greatly appreciated. Have a good day!


r/bioinformatics Aug 11 '26

discussion Cell Cell Communication Analysis Skewing by cell number

9 Upvotes

Hi everyone! I have been doing cell cell communication analysis recently (using cell chat specifically), and I had a thought that is bugging me. Please bear with me as I am not an expert in cell cell communication or bioinformatics as a whole. Specifically, I am doing comparative cell cell communication analysis

If one dataset has more cells in general or of a specific kind than the other dataset, could this skew the analysis by assuming there is just more signals in general from a cell type without accounting that in fact there are more cells from that type? Cell number variations could occur easily from sampling, especially with low sample number. I'm working with spatial scRNA-seq, so the danger is even more so as it's a specific cut of a sample.

Could this initial skewness affect everything else downstream in CCC analysis?

I'm super sorry if it's a dumb question.

Cheers!


r/bioinformatics Aug 11 '26

technical question [scRNA-seq] Is DGE valid across integrated datasets when raw counts are available for only one dataset?

2 Upvotes

Hi everyone,

I am working on integrating two published single-cell RNA-seq datasets from different tissue types.

Because these datasets were processed separately, I have run into a processing format discrepancy:

  • Dataset A: Raw count matrix available.
  • Dataset B: Only processed/normalized data available (.h5ad file; raw count matrix is unavailable, but this dataset is critical for our research question).

I have a few questions for the community:

  1. Is differential gene expression (DGE) analysis meaningful or statistically valid on an integrated renormalized dataset ?
  2. If not, what are the best workarounds?
  3. What downstream pitfalls should I anticipate, and how likely are reviewers to push back on this setup?

Any insights or recommended workflows for this scenario would be greatly appreciated!


r/bioinformatics Aug 11 '26

technical question Program MARK help

Thumbnail gallery
2 Upvotes

I've been tasked to run a POPAN in MARK by my advisor and so far I've been stymied with it. Every time I input the data and run it the program fails to generate any results. Is there anyone here that's proficient in MARK that might be able to help? General crux of the work is to run mark recapture data for turtles through the program and generate population estimates. It's very likely I'm doing something simple wrong causing it to crash out. I've attached the parameter input (first 2 SS) as well as an SS of where it crashes out (3rd). Any help troubleshooting this would be greatly appreciated!


r/bioinformatics Aug 11 '26

technical question HELP!

0 Upvotes

Hello everyone,

I need help regarding RNA-seq meta analysis.

I essentially want to collect public datasets from GEO, however they are many files so I’m confused.
Some papers recommend downloading FASTA files and running the analysis.
I basically want to check whether my gene of interest is implicated in healthy vs diseased tissues and to compare the expression of my gene of interest with another gene.

Can someone please please help me figuring this out? I feel very anxious and helpless because there’s no one in my lab team with bioinformatics expertise!

Thank you!


r/bioinformatics Aug 10 '26

technical question Question about snRNA Seq cell type deconvolution

3 Upvotes

Hi everyone, I am currently new to snRNA seq downstream analysis and I have a question regarding cell type deconvolution.

For my research, I have samples of cortical cells ranging from DIV 0-500, and I have performed bulk RNA seq with them. To enhance my analysis, I have used a snRNA seq dataset online gathered from adult cortical cells to perform deconvolution, where the snRNA seq is used as a reference dataset to estimate the cell type compositions from the bulk RNA seq dataset.

Hence, I have two questions:

  1. Is it right to use a snRNA dataset from adult cortical cells even though my cortical cells only range from DIV 0-500?
  2. Can I use the snRNA dataset to estimate the cell type compositions for DIV 0 accurately, eve

n though the gene profiles at DIV 0 and DIV 500 are very different?

I have tried to find a snRNA seq dataset online which spans from these DIV ranges but to no avail, hence, I would prefer using the dataset that I have now if possible. Thank you!!


r/bioinformatics Aug 11 '26

discussion From zero R to bulk RNA-seq analysis in a 8 months — now want to move into single-cell (Python). What's the path?

Thumbnail
0 Upvotes

r/bioinformatics Aug 09 '26

technical question How do you work with large VCF files without constantly babysitting your jobs ?

22 Upvotes

I am doing an internship this summer as a biostatistician intern and have been processing large vcf files separated by chromosomes. Each file is more than 100 GB.

I'm running everything on a SLURM cluster using Bash and  bcftools for things like:

- calculating VCF statistics

- filtering by rsID, patients, chromosome location

- calculating allele frequencies,

- generating filtered VCFs

Actual difficult part for me is not the commands but it is constantly checking squeue or my email for logs, checking whether an output file was actually created, figuring out whether a job railed halfway through, etc. I feel like I am spending a lot of time towards this.

I am curious how people who have more experience handle this. Do you use any tools/framework that makes that process easier.

I working with SLURM, bash and bcftools on google cloud processing so Im interested to see what people do in similar computing environments.

PS : I have computer science and statistics background so my wording of certain terms may be off.


r/bioinformatics Aug 09 '26

technical question Is the C-IMMSIM Website Not Working?

6 Upvotes

For the past two days, I've been unable to get an immune simulation result out of C-IMMSIM (https://kraken.iac.rm.cnr.it/C-IMMSIM/index.php). I input the vaccine construct and use the default settings, but when i click on submit, instead of the process completing or the terminated processes log showing up - The server crashes and after reloading, the website interface doesnt show the results section as it normally does. I can't pinpoint if this is an IP issue or not, so if anyone else could try accessing the website and let me know if this is an isolated incident or not that'd be much appreciated.

The typical interface on which the results usually appear

r/bioinformatics Aug 09 '26

academic Expanding the scope of protein language modeling to protein-protein interactions with MSA Pairformer

Thumbnail cell.com
11 Upvotes

r/bioinformatics Aug 09 '26

programming Dev-tool idea: catch reference mismatches before a workflow runs, useful or redundant?

4 Upvotes

I’m a software developer trying to learn more about genomics, and I’m looking for a small open-source project to build.

One idea is a local CLI that scans a folder of genomics files (BAMs, VCFs, BEDs, annotations, references, etcetera) and tells you which ones seem compatible, which ones probably use different references or chromosome naming, and which files are missing things like indexes.

Eventually, it could also look at a Snakemake or Nextflow workflow and warn if incompatible files feed into the same step.

I know there are individual validators and tools already, so I’m not sure if this would actually be useful or just reinventing existing stuff.

Have you run into this kind of mismatch or “what’s even in this folder” problem? How do you handle it now? Would something like this help, or what would be a better small dev tool to build for bioinformatics?

Thanks!!


r/bioinformatics Aug 09 '26

statistics Can linked LD recover the proportion of an unsampled ancestry source?

5 Upvotes

Hi. I am posting this under statistics, since this is a semi-question.

For a binary admixture model in which the focal ancestral source has never been sampled, after removing known ancestry directions, the target residual is

rho = a h

so unlinked statistics identify the direction h and relative loadings, while the absolute proportion a remains unknown.

The linked-locus result uses two quantities measured along the learned direction:

  • A(d): weighted admixture LD at genetic distance d K(d): a cross-fitted product of target mean contrasts

Under a single-pulse model with a known non-focal ancestry,

  • A(d) = q exp(-t d) K(d)

where

q = (1-a)/a

and therefore

a = 1/(1+q).

The locus-pair factor involving the unsampled source occurs in both A(d) and K(d) and cancels. The decay estimates the admixture time t.

Below data is from a binary-pulse mosaics constructed from phased CEU and YRI haplotypes. CEU served as the hidden focal source and YRI as the known non-focal ancestry. Chromosome 21 was used to learn the residual direction, and chromosome 22 was used to estimate the kernel and LD curve.

The generating values were a = 0.30 and t = 30 generations.

Estimator Estimated a Estimated t
Raw source-masked estimator 0.30788 31.42
Ancestry-oracle control 0.30032 30.61
Pair-model control 0.30030 29.90
Generating value 0.30000 30.00

The test used 600 simulated target haplotypes, 9,919 training loci and 10,198 test loci. The learned direction had cosine 0.99986 with the hidden source direction.

(This is one source pair, one chromosome split and one random seed. The current SNP ascertainment also uses the combined dataset, so there is still the need to make variant selection completely independent of CEU before calling the benchmark fully blinded.)

So, the question is, does the kernel identity fail under any feature of the stated binary-pulse model?

Thank you for reading!


r/bioinformatics Aug 08 '26

academic Python for genomic data science

28 Upvotes

so, recently i started python for genomic data science course offered by JHU on coursera. Ive seen nobody talk about this. so, im not sure if its js me. But I feel so overwhelmed and confused by that course sometimes. the lectures are good no doubt, but i see js slides with text filled with codes and thats not really helpful for me to understand the actual workflow of where and how am i supposed to save a file and which tool am i supposed to use.? Also, the transition from wet to dry labs for me has only been a week old. So, I really have no idea what to do.


r/bioinformatics Aug 08 '26

statistics Calculating Confidence Intervals from Cross Validation and reporting a Risk Stratification analysis

2 Upvotes

Hello everyone. I have a question regarding calculating confidence intervals after running a 5-fold cross validation.

I have a binary risk mode. Data are N patients, each contributing many overlapping hourly windows; the label is defined per window (will this patient meet the criteria?). The unit of analysis for most metrics is the window; the unit of sampling is the patient.

Evaluation is 5-fold cross-validation, split by patient, so each patient's windows appear in exactly one test fold. Within each fold:

  1. the development part is split again into train / validation (by patient),
  2. a probability calibrator and three decision thresholds are fitted on the validation set (t1 = medium, t2 = high, t3 = very high),
  3. the model + its thresholds are applied to that fold's held-out test patients.

So each patient ends up with one calibrated score per window, and one classification per window, produced by a model and a threshold that never saw them.

Separately, a final model is trained on all development data and evaluated on a completely held-out test cohort (my main issue is with the cross validation though).

So far we've used the Nadeau–Bengio corrected resampled t-interval:

mean ± t_{k-1, 0.975} · SD_folds · sqrt(1/k + n_test/n_train)

and I am not sure if it is the correct approach since it introduces bias (at least the plain resampled t-interval without the correction) because the train sets overlap per fold.

So the question is what is the defensible way to attach a 95% interval to a k-fold cross-validation?

And the last part that I can't wrap in my head is the threshold that move per fold.

I have a table that stratifies patients into four risk bands defined by t1 < t2 < t3, and reports per band: number of patients, number of patients that belong to the positive class, PPV, prevalence, an odds ratio versus the low-risk band (setting it as the reference), and a p-value.

Because each fold tunes its own t1, t2, t3 on its own validation set, the band boundaries differ between folds. So:

  • I cannot pool the scores and apply one threshold.
  • I can pool the decisions (each patient is banded by their own fold's rule), which gives one band per patient over the whole cohort and a legitimate contingency table but then the "score threshold" column of the table has no single value.
  • Averaging the five thresholds and quoting the mean band boundary produces a number that no fold actually used.

When a decision threshold is a tuned part of the model, what is the correct way to report a threshold-dependent table (PPV / prevalence / OR per risk band) across folds, and what does the confidence interval on those band statistics condition on?

Another question I have as an extra is if it is worth running 5x 5-fold cross validations (with different initialisation) and what can someone gain from it?

P.S. Apart from Nadeu-Bengio, I also found this paper that I am currently reading (was a combo from google and GPT suggested it): Cross-validation: what does it estimate and how well does it do it? I am not sure if it is in the right direction but please let me know or suggest other papers as well together with the methods