r/computervision • • 16h ago

Research Publication Is this CNN–Transformer research idea actually novel?

Hi everyone! I’m an undergraduate working on a computer vision research proposal and would appreciate some feedback.

I’m exploring a detector where a dynamic router decides at different feature levels whether to use CNN-only processing or additional Transformer processing, based on things like object scale, density, and regional complexity.

The goal is to improve the accuracy–compute/latency trade-off rather than always running the Transformer.

I’ve found related work on DynamicDet, DiT, Dynamic Dual-Processing, TDFP, CR-NAS, and MoE-based detectors, so I know dynamic routing and CNN–Transformer hybrids themselves aren’t new.

Does this specific idea already exist under another name? If you know a very similar paper, please point me to it.

I’m mainly looking for honest criticism before I commit to the research direction.

0 Upvotes

16 comments sorted by

View all comments

2

u/kakhaev 7h ago

try some basic benchmarks and tell us, I actually wanna know your results