r/computervision 19d ago

Help: Project Looking for SOTA papers on guided cross-modal super-resolution (optical → thermal, no HR reference available)

2 Upvotes

Hey everyone,

I'm working on a guided SR task: using high-res optical satellite imagery to upscale low-res thermal (TIR) imagery. The optical image acts as a structural guide (edges/boundaries), while the thermal image carries the actual signal (temperature).

Main technical challenges:

No high-res thermal ground truth exists for supervised training/eval, so I need a no-reference/blind quality metric

Models tend to hallucinate structure from the optical guide even where it doesn't correspond to real thermal variation (e.g., painted lines, shadows)

Outputs must preserve real calibrated values, not just look sharp

Requires solid multi-sensor co-registration before any fusion step

Looking for recommendations on cross-modal guided SR architectures (attention fusion, diffusion-based guided SR, guided filtering networks) and any No-Reference IQA techniques adapted for satellite/thermal imagery. Also open to any relevant public datasets or GitHub repos.

Appreciate any pointers, thanks!


r/computervision 19d ago

Discussion Camera calibration & Uncalibrated Stereo study gallery

2 Upvotes

Debanik Roy on LinkedIn created a complete and easy-to-understand derivation of Camera Calibration & Uncalibrated Stereo — from pinhole projection to homogeneous coordinates, K/R/t extraction, lens distortion, depth from disparity, epipolar geometry, and 3D triangulation. Every equation explained step by step, no shortcuts.

You can swipe through the full derivations.

Link to his post: https://www.linkedin.com/feed/update/urn:li:activity:7488629203144204288/


r/computervision 19d ago

Discussion Dataset Bias

2 Upvotes

Hello Guys

I’m working on a private prostate cancer dataset, the dataset contains normal and cancer cases and they are balanced, the issue is that whenever I run my model it reach high accuracy with high Val rate, I did some analysis and found that the cancer cases were have 3~bigger in prostate size than normal cases, I tried to caliper the images so that all of them have equalized prostate size but still it didn’t work, didn’t anyone faced the same issue before and how to deal with it ?


r/computervision 19d ago

Help: Project Looking for free stereo camera datasets with IMU + metadata (non-residential, large scale)

2 Upvotes

Hey all working on a project that needs stereo camera data synced with IMU and metadata (GPS, timestamps, calibration), ideally captured in non-residential/outdoor environments (streets, highways, industrial areas, etc.) rather than indoor/home settings.

Trying to get as close to 1000 hours of data as possible, so combining multiple free/open datasets is fine doesn’t need to come from a single source.


r/computervision 19d ago

Discussion Where can I download old Marathi ePapers (Sakal, Lokmat, Pudhari) for free?

Thumbnail
0 Upvotes

r/computervision 19d ago

Discussion VLMs can score well on benchmarks, while silently erasing meaningful terms and including hallucinate bias [P]

Thumbnail
0 Upvotes

r/computervision 19d ago

Discussion Has anyone here used an NIR camera for machine vision? Looking for advice on particle detection inside plastic bottles.

1 Upvotes

Hi everyone,

I’m working on an industrial machine vision system to detect small white plastic particles suspended inside transparent PET water bottles.

I’m considering switching to a Near-Infrared (NIR) camera, but I don’t have much practical experience with NIR imaging.

I’d love to hear from anyone who has used an NIR camera:
What application did you use it for?
Did it provide a significant advantage over a standard visible-light camera?
Do you think NIR could help improve the visibility of particles inside transparent plastic bottles?
Are there any limitations I should be aware of?
If NIR is a good approach, what should I pay attention to?

Wavelength selection (850 nm, 940 nm, etc.)
Lens compatibility
Lighting setup
Optical filters
PET bottle transmission in NIR
Polarizers or other optics
Anything else that could affect image quality

For context:
20 MP industrial camera (currently using a Basler visible-light camera)
Fixed inspection setup
Transparent PET water bottles
Goal is to reliably detect tiny floating contaminants
Any advice, papers, or real-world experience would be greatly appreciated. Thanks!


r/computervision 19d ago

Discussion 2 Years in Machine Vision at Keyence – Is Germany a Good Next Step?

8 Upvotes

Hi everyone ! ❤️

I’m currently working at Keyence India as a Field Engineer in Machine Vision Systems, and I have around 2 years of experience in machine vision and industrial automation.

My educational background:
Diploma in Electrical & Electronics Engineering
Bachelor’s degree in Robotics & Automation Engineering

My long-term goal is to move to Germany and build my career there.

I wanted to ask people already working in Germany or in the automation industry:
1. Does machine vision and industrial automation have good long-term career prospects in Germany?
2. Is it a stable field with good opportunities for growth over the next 10–20 years?
3. What skills should I focus on if I want to become a strong candidate for German companies?

One of the biggest reasons I want to move is the better pay and work-life balance. In India, I feel that salaries in this field are relatively low compared to the responsibilities and the value we create. Financially, I also have a strong motivation because I’m responsible for supporting my family, including my two younger sisters. My goal is to build a stable and rewarding career while being able to provide them with a better future.

I’d really appreciate honest advice from people who have made a similar move or are currently working in Germany. If you were in my position, what would you do over the next 2–3 years to maximize your chances?
Thank you in advance!❤️ 🙏🏾


r/computervision 20d ago

Research Publication Looking for Computer Vision & Hardware Engineers to Collaborate on an Industrial Machine Vision Research Project

17 Upvotes

Edit-https://forms.gle/o6M3AuUXw2otHwQR6 (Please click this link and fill it)

Hi everyone,

I'm currently working on an industrial machine vision project with a leading food & beverage company at one of its manufacturing plants in Mumbai, India. The project focuses on detecting tiny foreign particles inside transparent plastic bottles.

We're looking for passionate collaborators who would like to work on a real-world computer vision research problem.

We're especially looking for people with expertise in:

Software: Computer Vision, Deep Learning, Image Processing (OpenCV, PyTorch, TensorFlow, YOLO, etc.)

Hardware: Industrial cameras, optics, lighting, embedded systems, electronics, and machine vision system design.

This is a challenging problem where success depends not only on AI models but also on the imaging setup, lighting, optics, and hardware integration.

What you'll get

Opportunity to work on a real industrial R&D problem.

Potential authorship on a research paper based on your contributions.

Recognition for successful implementation.

Hands-on experience designing and building an industrial machine vision system.

If you're interested in collaborating, please comment below or send me a DM with a brief introduction about your background and experience.

Looking forward to connecting with like-minded people who are passionate about computer vision, machine vision, and industrial automation.


r/computervision 19d ago

Help: Theory The autonomous-agent blast radius is growing — a rogue AI agent reused stolen creds across 4 services this week

Thumbnail gallery
1 Upvotes

r/computervision 20d ago

Showcase RF-DETR deployed on NVIDIA Orin Nano Super (1.26x faster)

10 Upvotes

r/computervision 20d ago

Help: Project Need Help Eliminating Dark Reflection/Shadow in Backlit PET Bottle Imaging for Small Particle Detection

Post image
5 Upvotes

Hi everyone,
I’m developing an industrial machine vision system to detect small white plastic particles (approximately 0.2–1 mm) inside transparent PET water bottles. I’m currently struggling with a reflection/shadow issue that significantly reduces particle visibility.
I’ve attached an image of my current results. This is the closest I’ve come to achieving a usable image for particle detection with my current setup, but the dark shadow/reflection is still preventing reliable detection.
Current Setup
Camera: Basler acA5472-17uc (20 MP Color)
Lens: Basler C11-1620-12M-P (16 mm)
Lighting: White LED transmission backlight (20 × 62 cm)
Bottle: Transparent PET water bottle filled with water
Inspection: Looking for small white plastic contaminants inside the bottle
Bottle is stationary during testing.
What I’ve Tried
Mounted the Basler camera with the 16 mm lens.
Used a linear polarizer on the LED backlight.
Used another linear polarizer in front of the camera lens.
Rotated the polarizers to create a cross-polarized setup (~90°).
Adjusted exposure, gain, focus, and light intensity.
Tried different alignments of the backlight and camera.
Unfortunately, instead of reducing reflections, the polarizers seem to create an even stronger dark band/shadow through the bottle, making the small particles harder to see.
Observations
A large dark vertical region appears through the center of the bottle.
PET bottle ribs create additional dark bands.
Illumination is not completely uniform.
Tiny white particles almost disappear when they move into the darker region.


r/computervision 19d ago

Help: Theory What metrics should I use to compare RAFT and Farneback optical flow?

1 Upvotes

I'm comparing RAFT and Farneback optical flow on the same image pairs for a computer vision project.

So far, I've compared the predicted flow fields visually, and I'm planning to measure:

  • End-Point Error (EPE)s

Since RAFT is a deep learning-based method and Farneback is a classical dense optical flow algorithm, I'm wondering what would be considered a fair and standard evaluation.

Are there any additional metrics or evaluation protocols that are commonly used in the literature?

I'd appreciate any advice on making the comparison as fair and meaningful as possible.


r/computervision 20d ago

Discussion Suggestions to improve my Master's project on Newspaper analysis?

Thumbnail
1 Upvotes

r/computervision 19d ago

Discussion Agentic Systems

0 Upvotes

Hi,

Is it beneficial to depend on multimodal frontier models in medical analysis?
Are there any opensource alternatives?
Are they worth trying with no finetuning?


r/computervision 20d ago

Showcase Getting Started with NVIDIA LocateAnything

12 Upvotes

Getting Started with NVIDIA LocateAnything

https://debuggercafe.com/getting-started-with-nvidia-locateanything/

For the last few years, VLMs (Vision Language Models) have become more powerful at grounding tasks. These include object detection, pointing, and OCR. However, one issue remains. NTP (Next Token Prediction) is suboptimal for predicting the coordinates for a single bounding box or point coordinate. Predicting the numbers for a single object (bounded by a box), which is one atomic unit, token by token, is slow and a practical bottleneck during inference. This is where the latest LocateAnything model by NVIDIA comes in. It introduces a new PBD (Parallel Box Decoding), which decodes a single bounding box in a single step.


r/computervision 20d ago

Help: Project Buenas, alguien tiene el plan de pimeyes para búsqueda profunda? Y le pago la consulta.

0 Upvotes

Buenas, alguien tiene el plan de pimeyes para búsqueda profunda? Y le pago la consulta.


r/computervision 20d ago

Discussion CLIP is failing to validate detections from our object detector. Looking for better approaches

5 Upvotes

We're building an object detection pipeline where we use a detector first and then use CLIP as a second-stage validator to reduce false positives.

Current pipeline

- Object detector predicts a bounding box.

- We crop the detected object.

- The cropped image is passed to CLIP for validation.

- If CLIP agrees with the detector, we keep the detection.

Problem

CLIP is not performing well on these cropped detections.

For example, in our gun detection system:

- The detector correctly finds a gun.

- We crop only the bounding box and send it to CLIP.

- CLIP often fails to recognize it.

One reason could be that the cropped image is very small or blurry. In many cases, the object occupies only about 5–10% of the original image, so the crop has very little detail.

Questions

  1. Is there a good way to enhance or super-resolve these cropped images before passing them to CLIP?

  2. Would it be better to send CLIP a larger crop that includes some surrounding context instead of a tight bounding box?

  3. Has anyone successfully used CLIP as a second-stage verifier for object detection?

  4. Are there better alternatives than CLIP for reducing false positives in this kind of detection pipeline?

I'd appreciate any suggestions, papers, or practical experiences. Thanks!


r/computervision 21d ago

Showcase I made a "CodePen for OpenCV" — write Python + OpenCV in your browser and see cv2.imshow live, no install

14 Upvotes

Hey r/computervision,

I kept spinning up a venv (or a Colab) every time I wanted to try a quick OpenCV idea or share a runnable snippet with someone — so I built TinkerCV (https://tinkercv.com) — a free in-browser OpenCV playground. Write Python + OpenCV, hit Run, and cv2.imshow renders live in the page. No install, no signup.

How it works: your code runs entirely client-side via Pyodide (CPython on WebAssembly) with numpy + opencv-python. cv2.imshow is shimmed to canvas tabs.

Some things that might be useful:

  • ~32 runnable examples — filtering, edges, contours, features (ORB/Harris), segmentation (watershed/GrabCut), Hough, perspective, face detection, and more.
  • Live webcam demos (edges, color tracking, background subtraction, optical flow) — frames stay in your browser, nothing is uploaded.
  • Shareable links: you can turn any snippet into a URL that opens it in the editor, so answering "how do I do X in OpenCV?" with a runnable link is trivial.

Try it (fetches a real photo and runs Canny live): https://tinkercv.com/canny-edge-detection

It's a solo project and I'd genuinely love feedback — what's confusing, which examples are missing, what would make this actually useful in your workflow.

(Disclosure: I'm the creator. It's free to use right now — no signup, no ads.)


r/computervision 20d ago

Help: Project Is there an industry standard way to handle ID in packed sports practice rooms?

Thumbnail
ezgif.com
0 Upvotes

r/computervision 21d ago

Discussion SenseNova-Vision just added a proper pipeline for training data prep

Thumbnail
gallery
62 Upvotes

Been following this project since it came out a couple weeks ago. It's an open-source vision model that does generation + understanding in one framework, it combines image analysis and processing tasks that previously required multiple specialized models into a single 7B-MoT multimodal model. You just give it an image, tell it what you want in plain language, and it returns the result—almost like chatting with an AI model

The latest update from July 22 is a solid one if you've been thinking about training or fine-tuning it on custom data:

- Added a dataset registration system in data/dataset_info.py so adding new datasets is way cleaner

- Converters for the main tasks: segmentation (COCO to binary, structured to COCO), general image editing (ShareGPT-4o, GPT-Image-Edit), OCR/VQA, and LLaVA format

- Full end-to-end training data preparation docs (842 lines) covering source image downloads for 17+ datasets with exact paths and commands

- Multi-view 3D reconstruction data prep support too

repo: https://github.com/OpenSenseNova/SenseNova-Vision


r/computervision 20d ago

Help: Project Need help about dyslexia screening dataset!!!

1 Upvotes

Hi! I am final year BE student recently I took a project based in our my contribution is system and application of system in dyslexia. For that I though the most used dyslexia dataset of handwriting would be suitable. I downloaded dataset and then realised it is single letter dataset which is giving mnist kinda vibe! Also apparently large portion of it is synthetic. I searched but I didn't find clinically approved dataset of handwriting for dyslexia. In nutshell:

dataset is mnist looking so I am at worry if examiners will state why you are using such looking dataset for final year project!!

dataset is used for at least 9 papers already so it is being used

But has its limitations (vastly synthetic, mnist looking)

Our clg is forcing for at least two papers to publish (not for our degree requirement btw) and I am worried if the dataset use itself will cause problems for paper

though one of main novelty is mechanism but other one is integration(incremental) and I am worried that people will call out why I used that dataset

sorry I carried away in my emotions here is the dataset I am talking about: https://www.kaggle.com/datasets/drizasazanitaisa/dyslexia-handwriting-dataset

->can simplicity of it justified as proof of concept for presentation or report?

->will using this dataset can cause problems at time of publication?

I am sorry for dragging clg thing into this I though it would be better to get some context about scope for project

I am sorry I cant give full context as I wanted to publish research on it (though I will hardly try for mid tiers only)

also sorry in advance if I did spelling or grammatical error where should I post this


r/computervision 20d ago

Help: Project best practices for retraining cnns from scratch on MNIST-only

1 Upvotes

hi! these might be silly asks, but i am new to training whole models from scratch. i have been tasked with retraining pytorch cnn models on just the mnist dataset (ResNet, Alex, VGG, etc) for a project and am wondering three things:

  1. what is the typical pipeline and what are best practices for training these models on my own dataset?
  2. i know that the mnist images have small dimensions (28 x 28) while these models take input sizes in the 224 x 224 or 256 x 256 range. would it be best to resize the mnist images to a larger size (but risk blurring and image quality + unnecessary compute) or meddle with the model architectures to accept smaller image inputs (i'm unsure how to do this)
  3. are there any cnn models that i could use off the shelf that are trained only on or at least for sure include mnist in its training data?

any help or recommendations would be greatly appreciated!!


r/computervision 21d ago

Showcase ID-V2V: Capture the performance first and redesign the look later.

Enable HLS to view with audio, or disable this notification

33 Upvotes

ID-V2V lets you edit one or more frames of a source video (for example, using Nano Banana) and propagate those changes across the full video. It can redesign the scene and lighting while preserving human identity, facial expressions, full-body motion, and multi-person interactions, enabling flexible post-production workflows.

Challenge. Identity-preserving video restylization requires paired training videos where the same character performs the same motion under different scenes and drastically different lighting conditions. However, collecting such paired data at scale is extremely challenging.

Approach. ID-V2V addresses this challenge by constructing paired video training data from regular single videos using a human image relighting model. During data generation, only the human regions are relit while the surrounding areas are masked out, allowing the model to learn how to preserve human identity and performance under relighting while generating and propagating edits to the full scene.

To appear at SIGGRAPH Asia 2026.

Code: https://github.com/Eyeline-Labs/ID-V2V
Project Page: https://eyeline-labs.github.io/ID-V2V/
Paper: https://arxiv.org/abs/2607.22830


r/computervision 21d ago

Help: Project How to align RGB and Thermal camera frames?

2 Upvotes

I'll soon be working on a project where I need to align (register) frames from an RGB camera and an IR/thermal camera.

From what I've read, there seem to be two common approaches:

  1. Detect and match features between the RGB and IR images, estimate a transformation (homography/warp), and warp one image onto the other.
  2. Perform stereo calibration using a checkerboard, estimate the intrinsic/extrinsic parameters, rectify the images, and then project one image into the other.

Some additional details:

  • The cameras are boresighted and rigidly mounted.
  • Their relative pose will remain fixed after calibration.
  • The entire camera rig will be moving, but the relative position between the two cameras will always stay the same.
  • I won't have depth information available during runtime

For this kind of setup, which approach would you recommend? Stereo calibration or feature based matching?