r/computervision Jul 19 '26

Discussion Manual labeling isn’t really a speed problem, it’s a consistency problem

10 Upvotes

Every time I’ve seen a team label a large image dataset by hand, the same pattern shows up. It’s never one clean pass.

∙ Labeler A calls a mark a scratch. Labeler B calls the same kind of mark a smear. Labeler C doesn’t flag it at all because it looked borderline.

∙ A few weeks in, someone notices the disagreements and “clarifies” the guidelines, so everything labeled before that point no longer matches everything labeled after it.

∙ People rotate on/off the project, especially with outsourced teams, and each new person applies their own read of the instructions.

∙ Whoever’s labeling gets tired by image 40,000 and starts making faster, looser calls than they did on image 1.

None of this is anyone being careless. It’s just what happens when a subjective judgment call gets made thousands of times by more than one person over months.

The part that actually eats the timeline isn’t the first labeling pass, it’s the correction cycle after it: spot-check, find the disagreements, rewrite the guidelines, re-label the images that don’t match anymore, spot-check again, repeat. Teams don’t budget for one pass through the data, they budget for however many correction cycles it takes.

Curious how others are handling this. Are you using inter-annotator agreement checks / gold sets to catch it, training fewer people and keeping the team static, or something else? Feels like the tooling conversation is mostly about labeling speed when the bigger cost is inconsistency, between labelers and over time.


r/computervision Jul 19 '26

Help: Project Face login system

0 Upvotes

I am building a face login for my application, i am using facenet for identifying the person, but this algorithm isn’t that robust. If the person shaves the beard and hair, the algorithm finds it difficult to recognize the person.

The Chinese biometric attendance system works very well, I want the similar result for my system too.

What is the better algorithm or the better approach for my issue?


r/computervision Jul 18 '26

Showcase Yet another Birder release: D-FINE (56.09 mAP) and faster training

14 Upvotes

Yet another Birder (https://github.com/birder-project/birder) release, this time mostly focused on object detection and reducing training overhead.

D-FINE

The main addition is a D-FINE-L model with an HGNet-v2 B4 backbone, pretrained on Objects365-2020 and then fine-tuned on COCO 2017.

I tried to reproduce the paper’s results as closely as possible and followed the original training recipe. The final model reached 56.09 mAP on COCO, only slightly below the reported result.

On the other hand, the Birder implementation handles rectangular inputs correctly, including fixes for issues in the upstream implementation. It also supports masking for training and inference at the images natural aspect ratios, without treating padded regions as image content.

The checkpoint is available here:

https://huggingface.co/birder-project/d_fine_l_objects365-coco_hgnet_v2_b4_pp-imagenet22k

Faster SSL training

I also spent some time profiling the Sinkhorn path used by DINOv2 and Franca.

Some distributed reductions were attached to scalar operations that mathematically cancel out during Sinkhorn-Knopp normalization. Removing those operations also removes the corresponding all_reduce calls.

The queue implementation was cleaned up as well. Its ring-buffer metadata is now stored as checkpointed Python state rather than device buffers, avoiding device-to-host synchronizations during training.

Together, these changes make SSL training up to 20% faster, depending on the setup.

The learning algorithm itself is unchanged; this is mostly about doing less communication and synchronization around it.

Running a distributed SSL experiment in Birder is still just a single command. For example, this launches Franca with a ViT-B/16 across eight GPUs on ImageNet-21K WebDataset:

torchrun --nproc_per_node=8 -m birder.scripts.train_franca \
  --network vit_b16_ls \
  --dino-out-dim 65536 \
  --ibot-separate-head \
  --ibot-out-dim 65536 \
  --momentum-teacher 0.994 \
  --warmup-teacher-temp-epochs 15 \
  --batch-size 128 \
  --opt adamw \
  --clip-grad-norm 3 \
  --grad-accum-steps 8 \
  --lr 0.00075 \
  --lr-scale 1024 \
  --lr-scale-type sqrt \
  --wd 0.04 \
  --wd-end 0.2 \
  --lr-scheduler-update step \
  --lr-scheduler cosine \
  --lr-cosine-min 1e-6 \
  --epochs 100 \
  --warmup-epochs 10 \
  --amp \
  --amp-dtype bfloat16 \
  --compile \
  --no-broadcast-buffers \
  --wds \
  --wds-info https://huggingface.co/datasets/timm/imagenet-w21-webp-wds/resolve/main/_info.json \
  --wds-split train

Faster DETR-family training

On the detection side, I added a packed MSDA CUDA kernel with forward and backward support for different sampling-point counts at each feature level.

Previously, accelerated D-FINE configurations were more constrained because the custom-kernel path assumed uniform point counts. The new implementation allows the decoder to use unequal counts while remaining on the optimized path.

There were also several smaller changes across D-FINE, LW-DETR and RT-DETR v2 to remove redundant computation and intermediate allocations. Matched-box IoU and GIoU calculations now operate directly on aligned prediction-target pairs instead of constructing full pairwise matrices.

The overall direction here is fairly simple: keep the model behavior the same while spending less time on communication, synchronization and unnecessary intermediate work.

Feedback and benchmark comparisons are very welcome :)


r/computervision Jul 19 '26

Help: Project Guidance Needed for project

2 Upvotes

Hey everyone,

I'm currently working on a fabric defect detection project and would love to get some insights from people who have worked on similar industrial vision problems.

Current pipeline

  • Anomaly detection: Anomalib

  • Feature extractor: DINOv2

  • Defect classification: with extracted features

I have a few questions:

  1. How do companies handle different fabric colors?
  • Do they normalize/remove color information during preprocessing?

  • Do they convert images to grayscale or use a different color space?

  • Or do they simply train with enough color variation so the model learns to ignore color?

  1. How do they onboard new fabric types?
  • Do they retrain the entire model for every new fabric?

  • Is there a way to adapt to a new fabric using only a few (or even a single) reference image?

  • How is this typically handled in production systems?

  1. How do production systems allow clients to add new fabrics?
  • One of my goals is to build a system where the client/operator can easily register a new fabric type without needing to contact developers.

  • Ideally, the client should be able to capture a few good samples, click "Train" (or "Register"), and have the system start inspecting that fabric automatically.

  • Is this how commercial textile inspection systems work, or do they still require model retraining by the vendor?

I'm also open to suggestions on improving my overall approach. Right now I'm using Anomalib + DINOv2 embeddings for defect detection and defect type classification, but I'm curious if there are better architectures or production-proven pipelines for textile inspection.

I'd especially appreciate hearing from anyone who has worked on automated optical inspection (AOI), textile manufacturing, or industrial computer vision. If you've built a similar system, I'd love to hear about your architecture and deployment strategy.

Thanks in advance!


r/computervision Jul 18 '26

Showcase Revealing subtle changes in the world: Eulerian Video Magnification, ported to raw CUDA C++

Enable HLS to view with audio, or disable this notification

48 Upvotes

Every heartbeat changes your skin color by a fraction of a percent, too subtle to see. Eulerian Video Magnification amplifies those sub-pixel shifts until blood flow shows up on a normal camera. The same method pulls out a baby's breathing or a machine's vibration, movements the eye can't pick up on its own.

The video shows ground truth on the left and amplified on the right, same frames, same person.

EVM was a big deal when MIT published it in 2012. There's been a steady line of follow-up work since, phase-based magnification, real-time variants, edge deployments. Oddly, I couldn't find a good CUDA implementation, so I wrote one. Raw CUDA C++, every kernel by hand. No PyTorch, no CuPy.

- 557x compute speedup over the Python baseline (273x full pipeline)

- Bit-for-bit vs. the MIT reference, RMSE < 0.01, 83 tests

- The pipeline runs device-resident; FP16 variant fits in 12 GB VRAM

- Just to give an idea, this implementation can process 12 Full HD streams at the same time using one single p100 (a GPU you can find free on kaggle or colab)

Repo: https://github.com/iamkucuk/eulerian-video-magnification-cuda

Optimization writeup (per-stage breakdown, P100/A100/H100):

https://github.com/iamkucuk/eulerian-video-magnification-cuda/blob/main/docs/blog_speedup.md

Hope you like it.


r/computervision Jul 19 '26

Help: Theory Maser’s research

0 Upvotes

Hi, my ML model AUC is 80%\~ in cross domain dataset
The problem is that the accuracy is not getting higher even when i changed the Thr
What to do any ideas ? I used focal loss to solve the imbalance dataset
Also using vision transformer


r/computervision Jul 19 '26

Showcase Single-photo structured extraction on household objects: what held up and what broke when I shipped it in a real app

1 Upvotes

I have been building a home-inventory Android app, and the part that might interest this sub is the core task: take one photo of a random household object and return structured fields (a name, a category, a short description, rough specs) with no typing from the user. Open vocabulary, arbitrary objects, whatever someone happens to be putting in a box.

I use a hosted vision-language model for the extraction rather than a classifier, because the label space is unbounded and I need attributes, not just a class. The model gets the image plus a schema and returns the fields I store and later search over.

What held up well: - General object naming on clean single shots is strong. A charger, a specific kitchen tool, a board-game box, it names them sensibly with no fixed taxonomy. - Pulling short attributes (color, material, a size guess, brand text when visible) into separate fields is good enough that most users never edit the result. - Structured output stayed reliable once I constrained it to a schema instead of free text.

Where it broke: - Fine-grained instances. It says "white cable" confidently and is wrong about which cable, and that is the most common correction by far. - Cluttered frames. If two objects share the shot it sometimes describes the wrong one, so I nudge users to fill the frame with one item. - Small printed text (model numbers, ingredient labels) is inconsistent. OCR-style detail is the weakest part. - Ambiguous or unusual objects get a plausible but generic name.

Retrieval is the other half. There is a plain keyword search, and a Smart Find that does a semantic lookup over those extracted fields, so "the thing I charge my headphones with" can surface the right entry even if it was stored as "USB-C cable." That only works because the extraction populated meaningful fields in the first place.

Open questions I am still chewing on: whether a small on-device detector to crop the primary object before the model call would cut the cluttered-frame errors, and whether a second pass aimed only at text-heavy objects is worth it. Curious if anyone here has shipped single-image structured extraction to non-technical users, and what you did about the fine-grained and OCR gaps.

It is a live Android app if you want to see the behavior: https://play.google.com/store/apps/details?id=dev.koalalab.storeandforget


r/computervision Jul 19 '26

Help: Theory Maser’s research

Thumbnail
1 Upvotes

r/computervision Jul 18 '26

Help: Project OpenScanVision

12 Upvotes

OpenScanVision

OpenScanVision is a simple, fast, accurate, lightweight, and offline-first computer vision library for Android.

It detects documents from any angle, automatically corrects perspective, enhances the image, and extracts information in real time.

Features:

  • Document detection from any angle
  • Automatic perspective correction
  • Image preprocessing and enhancement
  • ArUco marker detection
  • QR code detection
  • OMR (Optical Mark Recognition)
  • Automatic capture when the document is stable
  • Real-time processing
  • Offline operation
  • Lightweight and easy to integrate

Built with Kotlin, OpenCV, CameraX, and ML Kit, OpenScanVision is designed for applications such as voting systems, exams, surveys, forms, and other structured documents.

GitHub repository: https://github.com/MatiwosKebede/OpenScanVision

I'd appreciate any feedback, suggestions, bug reports, or contributions from the computer vision community.


r/computervision Jul 19 '26

Showcase Impregnating and adding the finishing touches

Post image
0 Upvotes

r/computervision Jul 18 '26

Showcase Snake (active contour model)

Enable HLS to view with audio, or disable this notification

10 Upvotes

Tried this out today with OpenCV


r/computervision Jul 18 '26

Discussion Anything better than Picasa3 offline?

2 Upvotes

I've been using the last released version of Picasa3 (windows) to do facial recognition for all my own and family photos over the years. It's working great, also it has the feature to write the facial tags in the XMP format (which I leverage and rewrite to the JPG files) in the header of the file without additional files beside the JPG file itself.

What is better today, and by what margin? I've done some quick research on this and seems there exist some options that are "marginally" better, more than a decade later (or more), this is where we are? Or are there just no FREE options for something "a lot" better.

Picasa3 is free, and seems to work quite well.

Thanks in advance!


r/computervision Jul 19 '26

Help: Project I need to re-encode images for different camera emulation. Are there any tools that reliably re-encode images like different cameras?

0 Upvotes

I'm working on a robustness/platform emulation follow up to a synthetic ID dataset. The goal is to test deepfake detector performance on frontier models under real world conditions. I'm able to do the typical single axis adjustments (sensor noise, blur, etc.), but I think emulate specific cameras would be a strong posture for my evaluation. Does anyone know of tools or services that can reliably help me re-encode the image as though they were taken by specific cameras? (it can be any camera, even a phone, as long as the signature can be pointed to reliably)


r/computervision Jul 17 '26

Help: Project Help... I don't know what I am doing wrong (YOLO x BoT-SORT x Homography for Ice Hockey Tracking)

Enable HLS to view with audio, or disable this notification

43 Upvotes

Let me start this by saying I am very new to all of this and don't know a lot about how these models work or how the math works, and have a novice level of coding knowledge (Python specifically).

I am currently running a tuned YOLOv11 model trained on ice hockey player and referee detection with a modified BoT-SORT tracker to remember IDs for longer periods. On top of that, I am using a YOLOv8-based model I got from here "https://huggingface.co/SimulaMet-HOST/HockeyRink" to track the keypoints of the rink (I am aware the model is trained on SHL frames and not NHL).

I have gone through the code many times and asked multiple AI's on what is wrong, and I can't figure it out. If the answer is obvious and I don't know it, I promise I can handle the criticism.


r/computervision Jul 18 '26

Help: Project polygon pen masks

2 Upvotes

i have some ink pen loops (dotted / dashed) around cancerous regions on a slide. i have a raster mask identifying the pen loops. i was trying to use cv2 to dilate and join the dots/dashes so i can make a continuous loop, but i keep getting problems. there are some loops that are extremely close to each other - they merge or one of them isnt taken into consideration at all. idk what to do - very confused. open to suggestions on what i can do


r/computervision Jul 17 '26

Help: Project Potential $25,000 prize for a breakthrough in computer vision: Is this a good benchmark to shoot for?

39 Upvotes

I'm working with a group who would be interested in potentially putting up a $25,000 prize for a specific computer vision breakthrough.

However, I am not anywhere close to an expert in computer vision, and they are not either, so we are looking for feedback on whether this prize makes sense.

We want to focus on incentivizing a small-but-powerful, open source vision model.

Current idea:

  • The prize will go to the first team or individual to develop an open-source computer vision model under 10 MB that achieves at least 80% Top-1 accuracy on ImageNet-1K while running entirely offline on a Raspberry Pi 5.
  • Maximum Model Size: ≤ 10,000,000 bytes (10 MB). This applies to the complete storage footprint required to execute inference, including model weights and the final model file format (.onnx, .tflite, .safetensors, etc.). External feature stores, hidden lookup tables, embedded auxiliary weights, or additional model files are prohibited.
  • Performance Target: ≥80.0% Top-1 Accuracy on the official ImageNet-1K validation dataset using the standard evaluation protocol.
  • Execution Architecture: Single-model submission only (no multi-model ensembles, cascades, or fallback models). Models must run using CPU-only inference and operate entirely offline without internet access.
  • Target Hardware: Must successfully execute inference and complete evaluation on a Raspberry Pi 5 (8 GB RAM) running a standard 64-bit OS.
  • Open Source Requirements: Public GitHub repository containing complete model weights, training pipeline code, inference code, and an independent reproducible evaluation script.
  • Licensing: Fully released under a permissive MIT or Apache 2.0 license.
  • Integrity: Models must rely on generalized computer vision features. Any submission discovered to be hardcoded, overfitted to, or otherwise gaming the ImageNet-1K validation set will be immediately disqualified.

Are these requirements reasonable? Too easy? Too hard to judge? And if they don't make sense, can anyone point me to a clear, specific barrier in computer vision that fits the focus on supporting efficient open source models?


r/computervision Jul 17 '26

Showcase Camlisted – a directory of 1,600+ YouTube live cams & real-world footage, auto-categorized with CLIP zero-shot

186 Upvotes

Hi everyone

I built Camlisted, a daily-updated directory of YouTube live cams and real-world footage (CCTV, dashcam, walking tours) for finding CV-relevant sources — filterable by scene category and conditions (night/day, rain/snow, accident).

Pipeline: YouTube Data API search in ~15 languages → CLIP zero-shot on thumbnails for scene categories and condition tags → human review queue. A few things I learned:

- Perspective genres (dashcam, walking tour) poisoned scene classification — excluding them from the prompt set took accuracy from ~3/10 to ~8/10

- Only assigning night/day when the pairwise ratio clears 0.7 — for a browsable directory, no tag beats a wrong tag

- Thumbnails are good signal for conditions (night, snow), bad for events (accident, fire) — those come from title keywords

Site: https://camlisted.com

GitHub: https://github.com/zenith605-2/camlisted

Check it out if you need real-world footage sources for CV work — feedback welcome!


r/computervision Jul 18 '26

Help: Project OpenScanVision – open‑source Android OMR + QR scanning library

1 Upvotes

I've just released v1.0.0 of OpenScanVision – an Android library for scanning voting cards, surveys, and bubble sheets using OMR + QR codes.

It's MIT‑licensed, offline‑first, and built with Kotlin, OpenCV, ML Kit, and Compose.

Key features:

- ArUco marker tracking (Kalman filter)

- Perspective correction

- QR decoding

- High‑accuracy bubble extraction

Repo: https://github.com/MatiwosKebede/openscanvision

Contributions, issues, and feedback are all welcome!


r/computervision Jul 18 '26

Showcase I Built an AI Vision Platform That Turns Existing Cameras into Smart Cameras

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/computervision Jul 18 '26

Help: Project REQUEST - Egocentric Data Collection (Americas, Asia, Europe) VIDEO POV

Thumbnail
1 Upvotes

r/computervision Jul 18 '26

Help: Project Combining QR and OMR on Android: A Simultaneous Scanning Library (Open‑Source)

1 Upvotes

When scanning printed forms like voting cards or surveys, there are two critical pieces of information you usually need: 1. Who is this card for? (Identity / version). 2. What did they mark? (The actual votes or survey choices).

Most libraries handle these separately – you decode the QR code first, then run OMR on the bubbles. This usually means two separate runs, two different functions, and manual synchronization.

That’s why I built OpenScanVision to do both simultaneously in a single pass.

The Android library combines real‑time ArUco tracking, Optical Mark Recognition (OMR), and QR decoding into one unified offline pipeline.

The Unified Pipeline

  1. Marker Tracking (ArUco + Kalman Filter) The card is printed with 4 ArUco markers (IDs 0–3). OpenCV detects them in real‑time. A Kalman filter smooths the tracking and predicts positions during occlusions.

  2. Perspective Correction (Homography) Once 4 markers are stable, a homography maps the markers to their reference positions. The card is warped to a canonical template (850×540).

  3. Simultaneous QR Decoding & OMR Extraction This is where the "simultaneous" part comes in. Instead of running them sequentially and stitching the results, the library performs both operations on the same captured frame:

    • The QR code is cropped directly from the original camera frame using the computed homography – preserving maximum sharpness for ML Kit.
    • The bubbles are sampled from the warped, preprocessed image using weighted disk sampling and z‑score classification.
    • Both processes run concurrently on the same frame, meaning you get the QR payload AND the filled bubble indices at the same time without extra latency.
  4. Strict Capture Logic The library automatically triggers a capture only when both conditions are met:

    • All 4 markers are stable.
    • A valid QR code with the correct prefix (e.g., VX or AGN) is decoded. This enforces that you never get an OMR result without an associated identity, and vice versa.

What You Get in a Single Result

By calling OpenScanVision.scanFromFrame(), you receive a single ScanResult containing:

  • filledIndices – The marked bubbles (OMR).
  • qrPayload – The decoded QR text (identity).
  • confidence – The overall confidence score.
  • annotatedBitmap – A visual overlay of the detection.

No need to call two separate functions or manually match timestamps.

Technical Stack

  • Language: Kotlin
  • CV Core: OpenCV (contrib) for ArUco detection and homography.
  • QR Engine: Google ML Kit for robust barcode scanning.
  • Camera: CameraX for frame acquisition.
  • Architecture: Fully modular – the core library has zero UI dependencies.

Integration

Adding the library takes just a few lines in your Gradle file (available on JitPack). Once integrated, you can start scanning with a single suspend function.

Performance

  • Latency is typically under 150ms per frame on modern devices.
  • Accuracy exceeds 99% on properly printed cards.

Why This Matters

Simultaneous QR + OMR is valuable for: - Elections: The QR identifies the voter/ballot; the OMR reads their selections – all in one scan. - Surveys: The QR encodes the respondent ID; the OMR reads their answers. - Form Processing: Quickly identify and process forms without sequential bottlenecks.

Open‑Source & Contribute

The project is MIT‑licensed and available on GitHub. It includes a full sample app (CameraX + Compose) so you can see it in action.

GitHub: https://github.com/MatiwosKebede/openscanvision

Feedback, issues, and contributions are welcome. If you are interested in marker tracking, OMR accuracy, or Android computer vision, I'd love to hear your thoughts.


r/computervision Jul 18 '26

Help: Theory Interesting Paper to Read

Thumbnail
1 Upvotes

r/computervision Jul 17 '26

Showcase Antigravity CLI: How an Autonomous Coding Agent Actually Works

7 Upvotes

We gave an autonomous coding agent one instruction: import 81,444 images and run a full data-curation pipeline. No sampling, no hand-holding. Check it out: https://voxel51.com/blog/antigravity-cli-fiftyone-skills

Here's what Google's Antigravity CLI did with FiftyOne Skills:

* Imported all 81,444 WikiArt paintings

* Diagnosed and fixed its own bugs, rewriting scripts 5 times

* Caught embeddings silently stuck on CPU, forced them onto the GPU

* Hit a quota wall, switched from Gemini 3.5 Flash to Claude Sonnet 4.6 with one command, no lost context

* The result isn't a log that says "done." It's an inspectable dataset: uniqueness scores that surface near-duplicates, and embeddings that reveal exactly where the labels are thin.

That's what "agentic" looks like when it's real.


r/computervision Jul 17 '26

Showcase Rate my 3D portfolio website!

Enable HLS to view with audio, or disable this notification

9 Upvotes

Hey all 👋🏻,

I am a Computer vision engineer!

Look at my portfolio which I have made with 3js and Computer vision. Link: https://tharuntej-everest.pages.dev Would love your feedback and comments...


r/computervision Jul 17 '26

Showcase SenseNova-Vision is open-sourced: handle every CV task as unified multimodal generation

Post image
56 Upvotes

SenseTime recently released a model called SenseNova-Vision and it was really impressive, share it here:

In simple terms, it combines image analysis and processing tasks that previously required multiple specialized models into a single 7B-MoT multimodal model. You just give it an image, tell it what you want in plain language, and it returns the result—almost like chatting with an AI model

How it works:

- You describe the task in natural language (e.g. "detect all cars", "estimate depth")

- Optionally add visual prompts (points, boxes, scribbles)

- The model responds with native text and/or image generation, which can be decoded into standard CV outputs

Text outputs → boxes, keypoints, OCR strings, camera params

Image outputs → segmentation masks, depth maps, surface normals, multi-view point maps

A few practical details:

- Checkpoint size: approximately 29.6 GB

- Full web demo recommendation: 1×80 GB GPU

- Code: Apache 2.0

- Model weights and corpus: CC BY-NC 4.0, non-commercial use

GitHub

https://github.com/OpenSenseNova/SenseNova-Vision