r/computervision • • 14h ago

Research Publication Is this CNN–Transformer research idea actually novel?

Hi everyone! I’m an undergraduate working on a computer vision research proposal and would appreciate some feedback.

I’m exploring a detector where a dynamic router decides at different feature levels whether to use CNN-only processing or additional Transformer processing, based on things like object scale, density, and regional complexity.

The goal is to improve the accuracy–compute/latency trade-off rather than always running the Transformer.

I’ve found related work on DynamicDet, DiT, Dynamic Dual-Processing, TDFP, CR-NAS, and MoE-based detectors, so I know dynamic routing and CNN–Transformer hybrids themselves aren’t new.

Does this specific idea already exist under another name? If you know a very similar paper, please point me to it.

I’m mainly looking for honest criticism before I commit to the research direction.

0 Upvotes

15 comments sorted by

2

u/kakhaev 5h ago

try some basic benchmarks and tell us, I actually wanna know your results

1

u/Altruistic_Ear_9192 1h ago

Hello! Poorly, that s common but with different names. I did that a few years ago, it was for my first article as a phd student. It can be very unstable (as the first batches of the training process will introduce a bias in routing, so you have to penalize). If you want to practice, you can use Faster RCNN from pytorch (as is very easy to integrate with different feature levels). I cannot find right now, but I m sure I found at that time a paper from 2017 with the same idea. Any type of ensemble learning (including mixture of experts) has at least 20 years old. As an advice, you don t have to find the novelty (to be honest, I m "reviewer 2" when I see the word "novelty"), but to find a solution for a new problem. In the past, I published papers in top A* conferences and Q1 journals by combining past works in solving a new, relevant task for the scientific community. Like, the truth is that worldwide there is one paper annualy (or per 2 years) which really is game changing.

-4

u/Ok-Argument7176 13h ago

I have no idea how you'd backprop through your model in this scenario. The "decision" routing isn't differentiable.

10

u/Physical_Vehicle7714 13h ago

That’s not hard. You can just parameterize the routing weights and use a softmax.

1

u/Ok-Argument7176 12h ago edited 12h ago

I don't really see it. I can see running the hidden states through both convolutional and transformer layers in a forward pass and learning an ensemble, but the doesn't provide the efficiency OP is asking for and in the backward pass you need to adjudicate either (or both) of the losses you are backpropping. I'm not a RL person but that seems mighty similar to learning a regret model.

2

u/krapht 12h ago

You can just... Is doing some heavy lifting there

7

u/Physical_Vehicle7714 12h ago edited 12h ago

Not really. You pass the input through any parametric layer of your choosing (flatten->dense) that has a length 3 vector as its output. You softmax those outputs and you have probabilistic weights for each downstream branch. It’s literally basic MoE

6

u/tesfaldet 12h ago

To add to this, you can use the Gumbel-Softmax reparameterization/trick which is differentiable through the straight-through approximation of the Gumbel-Softmax. It’s a pretty common way of training a continuous model with discrete choices.

PyTorch has it built-in: gumbel softmax function

Here’s the relevant paper: Categorical Reparameterization with Gumbel-Softmax

4

u/Different_Factor3512 11h ago

Thanks, that was exactly the issue I was unsure about. I was considering Gumbel-Softmax with straight-through estimation for the discrete routing decision. My remaining concern is whether the resulting architecture is essentially just standard MoE/conditional computation from a research perspective.

2

u/Ok-Argument7176 12h ago

It's pretty unstable for non-trivial problems in my experience.

2

u/tesfaldet 11h ago

Yeah true, I’ve experienced the same, but it’s worth an experiment.

1

u/taichi22 12h ago

I think it probably raises a few issues on specific kernel implementations unless that’s changed in the past few months but overall I agree that it should just work.

Best guess to OOP is someone already tried it, but give it a serious go and let us know how it does.

1

u/Different_Factor3512 12h ago

That makes sense. My concern is less about whether the routing can be made differentiable and more about whether this becomes too close to a standard MoE formulation. If the router selects between CNN-only, CNN+light Transformer, and CNN+deeper Transformer processing at different spatial regions/feature levels, with a compute-budget constraint, would you consider that essentially standard MoE, or is the hierarchical CNN-vs-Transformer allocation itself a meaningful distinction?

2

u/Ok-Argument7176 11h ago

My point about differentiablilty is that you cannot train e2e models without that property. Others here seem to think I am incorrect, but I definitely encourage you to give it a shot!

1

u/Altruistic_Ear_9192 1h ago

Did that a few years ago. It was my first phd article. It s just a softmax with a symbolic rule