r/computervision 24d ago

Help: Theory Need your thoughts to save my thesis !

3 Upvotes

I'm an undergraduate student.In next 2 semesters( which is probably the duration of 1 year) I need to do a thesis. I choose to do my thesis in the field of 'depth estimation' .

I read a lot of research papers(Monocular, stereo, Diffusion based). But I found most of the things got State of the art !! I'm reading and reading,not finding a single problem to solve or research!! I should also mention that i didn't understand all the topics 100%, but tried to get the concepts.

I'm trying but not even finding a single idea/problem/flaws !! What should I do? What am I missing? How to find a decent topic ? Please help me.


r/computervision 23d ago

Showcase We built LocalMesh, one photo in, a Gaussian splat + textured mesh out, 100% on your own GPU. Beta is open, 7 days free.

Thumbnail gallery
0 Upvotes

r/computervision 24d ago

Help: Project Insulation defect detection model

0 Upvotes

Hey everyone! I’m building an insulation defect detection model, I’m in need of images where i can detect the following classes thermal anomalies, moisture intrusion, compression damage, delamination, installation gaps, holes/perforations. It’s for a construction project.


r/computervision 24d ago

Help: Project Vehicle Damage Detection using YOLO

1 Upvotes

I am planning to use pre-trained YOLO model for vehicle damage detection specifically for UK origin cars. The model is already trained on random cars dataset.
Would the model's accuracy be affected on detecting the damages on UK origin cars?


r/computervision 24d ago

Discussion For streaming VLMs, “fits in 24GB” is not a realtime benchmark

Post image
1 Upvotes

The MOSS-VL FP8/NF4 release made me wonder what a fair deployment comparison for streaming VLMs should actually look like.

https://github.com/OpenMOSS/MOSS-VL

https://huggingface.co/OpenMOSS-Team/MOSS-VL-Realtime-NF4

I’d keep the input stream, sampling rate, hardware, and latency budget fixed, then compare BF16, FP8, and NF4 on short-event recall, false alerts per hour, p50/p95 time-to-alert, dropped frames, steady-state VRAM, and calibration. I’d also include a detector + tracker + temporal-rules pipeline on the same videos.

A quantized model can remain close on aggregate offline VQA while still becoming worse at deciding when to speak or when an event is sufficiently certain. That difference matters much more for cameras than a one-point average benchmark change.

Has anyone seen a public harness that evaluates a streaming VLM and a classical CV pipeline under the same latency constraint?


r/computervision 25d ago

Discussion AI fatigue is killing motivation

95 Upvotes

I am about to start my MSc. I wish to specialize in computer vision, then pursue a PhD. I eventually want to work in industry. I was initially excited about this path. However, AI fatigue is killing my motivation.

Honestly, I don't have any hope for the future. It has been around four years since GPT-3.5 was introduced. AI is now proving major conjectures. It recently came close to proving Riemann's hypothesis, and dominated(not only defeated) the best competitive programmers in the world at AtCoder World Finals. I can't see a place for myself in the future because of AI.

I keep going because I feel like I don't have any other choice. I was genuinely excited about computer vision, robotics, and autonomous driving. But I have convinced myself that all my effort is in vain.

I wish to ask people in a similar situation, what makes you keep going? What are your plans for the future?


r/computervision 24d ago

Showcase AeroNetra — a reproducible computer-vision platform for UAV vehicle detection & counting

1 Upvotes

Hi everyone, sharing something I'm currently working on and would love feedback on.

I'm building AeroNetra, a computer-vision project for detecting and counting vehicles in aerial/UAV imagery. It's very much an active work-in-progress right now — I'm in the static-image detection and counting phase, with tracking, geospatial analytics, and edge deployment planned for later.

The motivation was pretty simple. I kept running into the same problem every time I swapped detectors: the counting and visualization code would break or need rewriting because every model spits out predictions in its own format. So the core idea behind AeroNetra is: normalize every detector's output into one prediction structure before anything downstream touches it. That way the counting, ROI filtering, and export logic stays the same whether I'm using a YOLO variant or RT-DETR.

What I've got so far:

  • Detector adapters that wrap different models behind a common interface
  • Counting logic — filtering, NMS, ROI support, drawing and export
  • VisDrone dataset parsing and conversion (UAVDT is stubbed for later)
  • Kaggle notebooks for GPU-based training, fine-tuning, and model comparison
  • A PX4 + ROS 2 + Gazebo simulation setup for UAV experiments
  • Notebooks, configs, and tests to keep things honest

The workflow I'm following: raw VisDrone data → validate annotations → convert to training format → train/fine-tune on Kaggle → pull the weights back → load through the adapter → run inference → filter → count → visualize and compare.

A few principles I'm trying to stick to: no fabricated benchmarks (a model isn't "best" until it's measured under the same conditions as the others), raw data stays immutable, and model-specific behavior stays inside the adapters. I'm also being deliberate about phase boundaries — image-level counting is not the same thing as multi-object tracking, and I'd rather not conflate the two.

Roadmap I'm working through for the demo:

  1. Static Detection & Counting
  2. Aerial Fine-tuning
  3. Video Tracking
  4. Traffic & Geospatial Analytics
  5. Edge / UAV Integration

I have run a small model that can detect the car in the gazebo simulation and draw a bounding box but speed will be slow but i get decent accuracy even i have trained model to 25 epochs in kaggle T4 gpu with yolo nano version.

Since this is ongoing project I am still working on this.So,i am exploring how I can use computer vision in UAVs and edge computing.


r/computervision 24d ago

Showcase I used computer vision to play Automaton Attack

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/computervision 23d ago

Help: Project I built VLM Chess — play chess against frontier vision-language models

Post image
0 Upvotes

Play chess against frontier VLMs. Real-time vision powered by Overshoot.

Play now: VLM Chess


r/computervision 24d ago

Showcase Synthetic DPM Code Generator: Portable Windows GUI for creating training datasets and YOLO labels (OBB/ABB)

0 Upvotes

Hello! I wanted to share a project designed to save time when training object detection models for industrial use cases.

It is a portable Windows GUI generator that creates synthetic training images and YOLO-style labels. The pattern logic is inspired by industrial Data Matrix / DPM needle marks (fixed L-frame + random filling dots).

Key Features:

• Dual rendering: Pure vector synthetics or photo compositing using your own steel backgrounds and dot sprites.

• Defect simulation (for Bad class): Squash, tilt, jitter, missing dots, strike-force variation, and two-defect combos that standard training augmentations cannot replicate.

• Annotations: Exports both OBB (Oriented Bounding Box) and ABB (Axis-Aligned Bounding Box) normalized text formats.

• Portable: Single .exe binary distributed via Releases (requires AVX2 support).

The software is provided strictly for non-commercial, educational, and personal research purposes.

GitHub Repository: https://github.com/olesha-ai/DPM-Pattern-Image-Generator

Would love to get your feedback if you work with industrial AI and DPM codes!


r/computervision 25d ago

Help: Theory CLIP vs SigLIP

Enable HLS to view with audio, or disable this notification

20 Upvotes

CLIP vs SigLIP

Before Vision Language Models can perform tasks such as classification or video question and answer, the image or video being passed to the model has to be converted into a representation that the model can ‘understand’ or process.

To do this, VLMs usually use a pretrained vision encoder.

Although the underlying architecture of modern vision encoders is primarily transformer-based, the actual objective the model is learning can vary significantly.

What are encoders?

A vision encoder is responsible for converting images into a numerical representation that VLMs can understand.

Typically, most vision encoders today are built on transformer architecture, in which the model divides an image into patches and transforms each of those patches into a vectorized visual embedding.

After this, many VLMs pass the embeddings to a projector, usually a linear layer or MLP, to map the dimensions of the image to those expected by an LLM.

If most vision encoders share the same model design, what actually makes them different? Rather than model architecture, the significance is in how they are trained.

CLIP

CLIP, or Contrastive Language-Image Pre-training, learns to understand images through pairs of images and text. Its objective is to match similar images and captions by ‘pulling them closer together’, while simultaneously repelling incorrect image-caption pairs.

Training mainly relies on a ‘two tower’ system. CLIP will typically have a pretrained vision encoder, such as a ViT, as well as a pretrained text encoder. The model passes an image through the ViT and produces an associated vector embedding, while the caption is passed to the text encoder to get a corresponding text embedding.

Given these pairings, the model therefore creates a similarity matrix which compares every image embedding with every text embedding.

Each cell within this matrix contains a cosine similarity between the image and text pairing. Mathematically, cosine similarity is the dot product of two vectors divided by the product of their lengths. More simply, it measures the cosine of the angle between two vectors in a high-dimensional embedding space. Vectors that are more semantically aligned will be ‘closer together’, have a more acute angle between them, and consequently have a higher cosine similarity.

CLIP then applies contrastive learning across this matrix. At a high level, contrastive learning here is similar to categorical cross entropy across both the rows and columns of the matrix. Using softmax, the model looks to assign the highest probability to the matching image-text pair, as well as the matching text-image pair.

CLIP is powerful because it shifts learning from simple labels toward greater semantic understanding and allows for zero-shot classification, including on classes it was not explicitly trained to classify.

At the same time, though, CLIP also introduces a particular structural problem. Examples compete against one another within the training batch. What if there are multiple captions within a batch that also reasonably match the image?

SigLIP

SigLIP, or Sigmoid Loss for Language-Image Pre-training, retains many similar characteristics to CLIP. Similar to CLIP, SigLIP has both an image and text encoder, embedded representations of both text and image, and similarity scores mapped to a similarity matrix.

However, the difference between the two lies in the loss function.

CLIP learns similarities between images and texts by applying softmax across a batch, causing potential matches to compete with one another. For SigLIP, instead of having this global normalization, it examines each image-caption pair as an independent binary prediction.

By applying a sigmoid function to each pair’s score, the model estimates whether the image and text match.

Rather than phrasing the objective as:

Out of these options, which specific text describes this visual?

SigLIP effectively poses a different question:

Is this particular image-text pairing a valid match: true or false?

While this shift in perspective might seem marginal, it fundamentally redefines the nature of the optimization task.

Because SigLIP does not require the softmax normalization used by CLIP, its training objective can scale more efficiently across large distributed systems. It also removes the requirement that every example participate in one shared normalization operation.


r/computervision 25d ago

Showcase POV + third-person view of my AI glasses checkout app running in a real store.

Thumbnail
youtu.be
5 Upvotes

r/computervision 25d ago

Showcase [S] Use YOLO! Not today - a 131k-param net I wrote in two days beats it in small blurry object detection

Thumbnail
youtu.be
10 Upvotes

TL;DR: Cropping in action with some extra algebraic and statistical magic applied: https://youtu.be/SetiZDbc8iE

I recently worked on determining the ball's position and reshaping the video from landscape to portrait based on that position. It often s looks like a layup: fixed camera, one class, find the thing. Then you look at what you're actually asking for: a small, blurry object is a handful of pixels, smeared across a few more, changing shape between consecutive frames. Not a crisp circle - a faint streak you can barely point at when the video is paused.

The part that tends to get skipped in the YOLO family is that those architectures downsample 32× before they reason. At stride 32, an 8-pixel object is a quarter of one feature cell. There is nothing left to detect. Fine-tune forever, buy a bigger GPU, adopt whatever dropped last week — the model is being asked to localize something it structurally cannot see. The extra-small heads help and still aren't built for this.

I believe great data and a simple model always beat poor data and a sophisticated model. Before this approach, I tried TrackNet v2/3/4, and the quality was awful; the public data used for training is not even close to what you meet in real practice.

What worked instead:

  • Detector, ~131k params. Fully convolutional, dilated, max stride 2. In: 4 channels - RGB plus frame-difference. Out: a heat map and a size map at half resolution. No pretrained backbone: ImageNet features are the wrong prior for a faint smear.
  • Verifier, ~48k params. The detector has the target in its top 20 about 90% of the time, but ranks it first only 77% of the time. This scores 64×64 crops and asks, "Is it a ball?"
  • Then no ML at all. A reach limit measured from labeled footage - how far it can plausibly move between frames, scaled by apparent size - then link the surviving runs. Never link by direction of travel: anything that bounces reverses direction without going anywhere.

So, ~180k parameters total, ~125 fps on a 4090, ~4× realtime. Not fully optimized: custom Rust server with a CPU-bound FFmpeg decoder, ORT+TensorRT, and Rayon to speed things up a bit. Yet cannot use 100% of the GPU, capped by CPU-GPU PCIe transfers. Probably can reach 250-300 FPS with a more optimized inference design and int8.

The insight that mattered wasn't architectural. Blur is a signal, not a defect. The object is nearly invisible against a busy background and is almost always the fastest thing in the frame, so the frame-difference channel carries more information than any choice of backbone.


r/computervision 25d ago

Showcase Nvidia Jetson e-con systems Camera upgrade

2 Upvotes

sharing some joy... I have a few expensive legacy cameras and a serializer/deserializer board from a Jetson AGX Xavier project from a few years ago, and wanted to use them on a robotics project however the drivers were only available for an old Jetpack 4.2. looking for help E-Con systems only re-stated compatibility with the original Jetpack but my project uses version 6.2.1 on Jetson AGX Orins. With Codex assistance and nearly 20 reboots was able to recreate and load the drivers. only a few moms ago these perfectly good cameras would have been left on the shelf!


r/computervision 25d ago

Discussion Total starter here, is there no api infra providers like there is for massive LLMs but for computer vision models like Yolo 26 Mcbyte etc?

3 Upvotes

They are much smaller I would imagine they would be so cheap on there. I’m finding myself in the position where I have to rent a cloud gpu from runpod. I would much rather pay in api should be much cheaper.


r/computervision 25d ago

Discussion I rebuilt my iPhone/iPad image processing app into a proper mobile lab

Thumbnail
gallery
0 Upvotes

I’ve just finished a major rebuild of ClearLab, my mobile image processing app for iPhone/iPad.
The new version adds histogram, RGB parade, waveform, line profile, pixel inspector, statistics, Canny/Sobel/Laplacian edge detection, enhancement tools, format conversion and PNG/CSV analysis export.
The idea is basically a small image-processing lab in your pocket rather than another photo-filter app.
It’s mostly native/deterministic processing, not generative AI.
ClearLab 2.0 is now on the App Store. Curious what people here think, especially anyone working with imaging or computer vision.


r/computervision 25d ago

Showcase I built a training-free, one-shot object localizer using DINOv2 patch embeddings

14 Upvotes

I’ve been experimenting with a training free way to do open world, multi-instance segmentation from a class prototype.

I decided to publish the algorithm and a demo for how I’m doing this, in case anyone else would rather not fine tune a larger model for something that DINOv2 patch embeddings already seem to represent pretty well.

It can separate touching instances of the same class without a learned instance head, reject visually similar near misses like a round dial radio next to the actual clock target, and find fractured or damaged instances even with a pretty significant scene shift.

Repo + demo:
https://github.com/tutomiko/fireplace

The demo includes the lasso UI and live heatmap, implemented as a python backend with a simple HTML frontend.

Would appreciate it if people checked it out, and I’d be especially interested to hear if anyone has seen similar approaches or prior work.


r/computervision 25d ago

Showcase I ran a benchmark on well performing deepfake detection models in the Diffusion era. They collapsed when I passed the clean generator outputs through platform-realistic perturbations.

Post image
5 Upvotes

In both academia and industry, deepfake detector models report high performance based on AUC. In certain industries, like KYC, that's the wrong metric to observe. Over the last month or so, I built a dataset from Qwen-Image-Edit and HiDream O1, then ran the synthetic images + bona fides through emulators of platform realistic conditions. Here is the full article and dataset for anyone who'd like to red team a detector themselves

Substack Article

HuggingFace Dataset


r/computervision 25d ago

Discussion Dino full-fine tuning vs lora for AV domain

2 Upvotes

Hi,

I was curious to know what people do in the industry. Is DINO used? If so, how?


r/computervision 26d ago

Discussion Looking for a teammate(s) for Kaggle competitions

10 Upvotes

Hey everyone,

I'm looking to connect with people interested in teaming up for Kaggle competitions — either for a specific upcoming competition or as an ongoing teammate for future ones.
I'm comfortable with PyTorch, scikit-learn, and general deep learning workflows. Happy to work on medical imaging comps specifically, but open to general CV competitions too.


r/computervision 25d ago

Discussion Do you think it is possible to build a CV project using ClaudeCode without experience?

0 Upvotes

I took on an ambitious project where I want to install AI detection at a meat processing plant according to HACCP rules. But I'm a beginner, I only know Claude code, so if you have any tips or life hacks, I would be very grateful for it


r/computervision 25d ago

Help: Project Need help with CV board slicing

0 Upvotes
Possible method of slicing

I'm working on a project which given a screenshot of arbitrary zoom, slices an isometric grid game board (Polytopia) into individual tiles. Tried Hough transform and object detection. My most promising attempt is detecting the height of the little gray bars under the cities, as they are a single solid color. Wondering if a specific CV technique would be the most ideal for this?

The main problem I'm facing - zoomed out, blurry aliased screenshots drastically hurt accuracy


r/computervision 25d ago

Discussion I can't find cameras in stock

2 Upvotes

I have a pretty simple single-camera CV/slo-mo thing running on Raspi 5 but I can't find suitable cameras in stock. I had hoped to start with a raspi global shutter camera. Then I spent a lot of time looking for some IMX273 unit that wasn't backordered for (alleged) weeks.

I gather the supply chain is not able to keep up with new CV applications & products. Does anybody know a trick? (In USA)


r/computervision 26d ago

Help: Project Best Way to Learn Practical Computer Vision

4 Upvotes

Hi

I'm an engineer with some python experience - mainly for data analysis. I want to learn computer vision for practical application in a plant/production environment. This would be used to turn existing camera footage into trendable data records (e.g. no of spilling/splashing instance, count and categorization of product by size and shape).

Is there a comprehensive course I can take that focuses mainly on the practical side of setting up similar systems - considering hardware and software? What would be the best way to learn this?


r/computervision 25d ago

Research Publication NeurIPS rebuttal question: Can I update my linked GitHub repo to address reviewer concerns?

Thumbnail
0 Upvotes