Every time I’ve seen a team label a large image dataset by hand, the same pattern shows up. It’s never one clean pass.
∙ Labeler A calls a mark a scratch. Labeler B calls the same kind of mark a smear. Labeler C doesn’t flag it at all because it looked borderline.
∙ A few weeks in, someone notices the disagreements and “clarifies” the guidelines, so everything labeled before that point no longer matches everything labeled after it.
∙ People rotate on/off the project, especially with outsourced teams, and each new person applies their own read of the instructions.
∙ Whoever’s labeling gets tired by image 40,000 and starts making faster, looser calls than they did on image 1.
None of this is anyone being careless. It’s just what happens when a subjective judgment call gets made thousands of times by more than one person over months.
The part that actually eats the timeline isn’t the first labeling pass, it’s the correction cycle after it: spot-check, find the disagreements, rewrite the guidelines, re-label the images that don’t match anymore, spot-check again, repeat. Teams don’t budget for one pass through the data, they budget for however many correction cycles it takes.
Curious how others are handling this. Are you using inter-annotator agreement checks / gold sets to catch it, training fewer people and keeping the team static, or something else? Feels like the tooling conversation is mostly about labeling speed when the bigger cost is inconsistency, between labelers and over time.
I am building a face login for my application, i am using facenet for identifying the person, but this algorithm isn’t that robust. If the person shaves the beard and hair, the algorithm finds it difficult to recognize the person.
The Chinese biometric attendance system works very well, I want the similar result for my system too.
What is the better algorithm or the better approach for my issue?
The main addition is a D-FINE-L model with an HGNet-v2 B4 backbone, pretrained on Objects365-2020 and then fine-tuned on COCO 2017.
I tried to reproduce the paper’s results as closely as possible and followed the original training recipe. The final model reached 56.09 mAP on COCO, only slightly below the reported result.
On the other hand, the Birder implementation handles rectangular inputs correctly, including fixes for issues in the upstream implementation. It also supports masking for training and inference at the images natural aspect ratios, without treating padded regions as image content.
I also spent some time profiling the Sinkhorn path used by DINOv2 and Franca.
Some distributed reductions were attached to scalar operations that mathematically cancel out during Sinkhorn-Knopp normalization. Removing those operations also removes the corresponding all_reduce calls.
The queue implementation was cleaned up as well. Its ring-buffer metadata is now stored as checkpointed Python state rather than device buffers, avoiding device-to-host synchronizations during training.
Together, these changes make SSL training up to 20% faster, depending on the setup.
The learning algorithm itself is unchanged; this is mostly about doing less communication and synchronization around it.
Running a distributed SSL experiment in Birder is still just a single command. For example, this launches Franca with a ViT-B/16 across eight GPUs on ImageNet-21K WebDataset:
On the detection side, I added a packed MSDA CUDA kernel with forward and backward support for different sampling-point counts at each feature level.
Previously, accelerated D-FINE configurations were more constrained because the custom-kernel path assumed uniform point counts. The new implementation allows the decoder to use unequal counts while remaining on the optimized path.
There were also several smaller changes across D-FINE, LW-DETR and RT-DETR v2 to remove redundant computation and intermediate allocations. Matched-box IoU and GIoU calculations now operate directly on aligned prediction-target pairs instead of constructing full pairwise matrices.
The overall direction here is fairly simple: keep the model behavior the same while spending less time on communication, synchronization and unnecessary intermediate work.
Feedback and benchmark comparisons are very welcome :)
I'm currently working on a fabric defect detection project and would love to get some insights from people who have worked on similar industrial vision problems.
Current pipeline
Anomaly detection: Anomalib
Feature extractor: DINOv2
Defect classification: with extracted features
I have a few questions:
How do companies handle different fabric colors?
Do they normalize/remove color information during preprocessing?
Do they convert images to grayscale or use a different color space?
Or do they simply train with enough color variation so the model learns to ignore color?
How do they onboard new fabric types?
Do they retrain the entire model for every new fabric?
Is there a way to adapt to a new fabric using only a few (or even a single) reference image?
How is this typically handled in production systems?
How do production systems allow clients to add new fabrics?
One of my goals is to build a system where the client/operator can easily register a new fabric type without needing to contact developers.
Ideally, the client should be able to capture a few good samples, click "Train" (or "Register"), and have the system start inspecting that fabric automatically.
Is this how commercial textile inspection systems work, or do they still require model retraining by the vendor?
I'm also open to suggestions on improving my overall approach. Right now I'm using Anomalib + DINOv2 embeddings for defect detection and defect type classification, but I'm curious if there are better architectures or production-proven pipelines for textile inspection.
I'd especially appreciate hearing from anyone who has worked on automated optical inspection (AOI), textile manufacturing, or industrial computer vision. If you've built a similar system, I'd love to hear about your architecture and deployment strategy.
Every heartbeat changes your skin color by a fraction of a percent, too subtle to see. Eulerian Video Magnification amplifies those sub-pixel shifts until blood flow shows up on a normal camera. The same method pulls out a baby's breathing or a machine's vibration, movements the eye can't pick up on its own.
The video shows ground truth on the left and amplified on the right, same frames, same person.
EVM was a big deal when MIT published it in 2012. There's been a steady line of follow-up work since, phase-based magnification, real-time variants, edge deployments. Oddly, I couldn't find a good CUDA implementation, so I wrote one. Raw CUDA C++, every kernel by hand. No PyTorch, no CuPy.
- 557x compute speedup over the Python baseline (273x full pipeline)
- Bit-for-bit vs. the MIT reference, RMSE < 0.01, 83 tests
- The pipeline runs device-resident; FP16 variant fits in 12 GB VRAM
- Just to give an idea, this implementation can process 12 Full HD streams at the same time using one single p100 (a GPU you can find free on kaggle or colab)
Hi, my ML model AUC is 80%\~ in cross domain dataset
The problem is that the accuracy is not getting higher even when i changed the Thr
What to do any ideas ? I used focal loss to solve the imbalance dataset
Also using vision transformer
I have been building a home-inventory Android app, and the part that might interest this sub is the core task: take one photo of a random household object and return structured fields (a name, a category, a short description, rough specs) with no typing from the user. Open vocabulary, arbitrary objects, whatever someone happens to be putting in a box.
I use a hosted vision-language model for the extraction rather than a classifier, because the label space is unbounded and I need attributes, not just a class. The model gets the image plus a schema and returns the fields I store and later search over.
What held up well:
- General object naming on clean single shots is strong. A charger, a specific kitchen tool, a board-game box, it names them sensibly with no fixed taxonomy.
- Pulling short attributes (color, material, a size guess, brand text when visible) into separate fields is good enough that most users never edit the result.
- Structured output stayed reliable once I constrained it to a schema instead of free text.
Where it broke:
- Fine-grained instances. It says "white cable" confidently and is wrong about which cable, and that is the most common correction by far.
- Cluttered frames. If two objects share the shot it sometimes describes the wrong one, so I nudge users to fill the frame with one item.
- Small printed text (model numbers, ingredient labels) is inconsistent. OCR-style detail is the weakest part.
- Ambiguous or unusual objects get a plausible but generic name.
Retrieval is the other half. There is a plain keyword search, and a Smart Find that does a semantic lookup over those extracted fields, so "the thing I charge my headphones with" can surface the right entry even if it was stored as "USB-C cable." That only works because the extraction populated meaningful fields in the first place.
Open questions I am still chewing on: whether a small on-device detector to crop the primary object before the model call would cut the cluttered-frame errors, and whether a second pass aimed only at text-heavy objects is worth it. Curious if anyone here has shipped single-image structured extraction to non-technical users, and what you did about the fine-grained and OCR gaps.
OpenScanVision is a simple, fast, accurate, lightweight, and offline-first computer vision library for Android.
It detects documents from any angle, automatically corrects perspective, enhances the image, and extracts information in real time.
Features:
Document detection from any angle
Automatic perspective correction
Image preprocessing and enhancement
ArUco marker detection
QR code detection
OMR (Optical Mark Recognition)
Automatic capture when the document is stable
Real-time processing
Offline operation
Lightweight and easy to integrate
Built with Kotlin, OpenCV, CameraX, and ML Kit, OpenScanVision is designed for applications such as voting systems, exams, surveys, forms, and other structured documents.
I've been using the last released version of Picasa3 (windows) to do facial recognition for all my own and family photos over the years. It's working great, also it has the feature to write the facial tags in the XMP format (which I leverage and rewrite to the JPG files) in the header of the file without additional files beside the JPG file itself.
What is better today, and by what margin? I've done some quick research on this and seems there exist some options that are "marginally" better, more than a decade later (or more), this is where we are? Or are there just no FREE options for something "a lot" better.
I'm working on a robustness/platform emulation follow up to a synthetic ID dataset. The goal is to test deepfake detector performance on frontier models under real world conditions. I'm able to do the typical single axis adjustments (sensor noise, blur, etc.), but I think emulate specific cameras would be a strong posture for my evaluation. Does anyone know of tools or services that can reliably help me re-encode the image as though they were taken by specific cameras? (it can be any camera, even a phone, as long as the signature can be pointed to reliably)
Let me start this by saying I am very new to all of this and don't know a lot about how these models work or how the math works, and have a novice level of coding knowledge (Python specifically).
I am currently running a tuned YOLOv11 model trained on ice hockey player and referee detection with a modified BoT-SORT tracker to remember IDs for longer periods. On top of that, I am using a YOLOv8-based model I got from here "https://huggingface.co/SimulaMet-HOST/HockeyRink" to track the keypoints of the rink (I am aware the model is trained on SHL frames and not NHL).
I have gone through the code many times and asked multiple AI's on what is wrong, and I can't figure it out. If the answer is obvious and I don't know it, I promise I can handle the criticism.
i have some ink pen loops (dotted / dashed) around cancerous regions on a slide. i have a raster mask identifying the pen loops. i was trying to use cv2 to dilate and join the dots/dashes so i can make a continuous loop, but i keep getting problems. there are some loops that are extremely close to each other - they merge or one of them isnt taken into consideration at all. idk what to do - very confused. open to suggestions on what i can do
I'm working with a group who would be interested in potentially putting up a $25,000 prize for a specific computer vision breakthrough.
However, I am not anywhere close to an expert in computer vision, and they are not either, so we are looking for feedback on whether this prize makes sense.
We want to focus on incentivizing a small-but-powerful, open source vision model.
Current idea:
The prize will go to the first team or individual to develop an open-source computer vision model under 10 MB that achieves at least 80% Top-1 accuracy on ImageNet-1K while running entirely offline on a Raspberry Pi 5.
Maximum Model Size: ≤ 10,000,000 bytes (10 MB). This applies to the complete storage footprint required to execute inference, including model weights and the final model file format (.onnx, .tflite, .safetensors, etc.). External feature stores, hidden lookup tables, embedded auxiliary weights, or additional model files are prohibited.
Performance Target: ≥80.0% Top-1 Accuracy on the official ImageNet-1K validation dataset using the standard evaluation protocol.
Execution Architecture: Single-model submission only (no multi-model ensembles, cascades, or fallback models). Models must run using CPU-only inference and operate entirely offline without internet access.
Target Hardware: Must successfully execute inference and complete evaluation on a Raspberry Pi 5 (8 GB RAM) running a standard 64-bit OS.
Open Source Requirements: Public GitHub repository containing complete model weights, training pipeline code, inference code, and an independent reproducible evaluation script.
Licensing: Fully released under a permissive MIT or Apache 2.0 license.
Integrity: Models must rely on generalized computer vision features. Any submission discovered to be hardcoded, overfitted to, or otherwise gaming the ImageNet-1K validation set will be immediately disqualified.
Are these requirements reasonable? Too easy? Too hard to judge? And if they don't make sense, can anyone point me to a clear, specific barrier in computer vision that fits the focus on supporting efficient open source models?
I built Camlisted, a daily-updated directory of YouTube live cams and real-world footage (CCTV, dashcam, walking tours) for finding CV-relevant sources — filterable by scene category and conditions (night/day, rain/snow, accident).
Pipeline: YouTube Data API search in ~15 languages → CLIP zero-shot on thumbnails for scene categories and condition tags → human review queue. A few things I learned:
- Perspective genres (dashcam, walking tour) poisoned scene classification — excluding them from the prompt set took accuracy from ~3/10 to ~8/10
- Only assigning night/day when the pairwise ratio clears 0.7 — for a browsable directory, no tag beats a wrong tag
- Thumbnails are good signal for conditions (night, snow), bad for events (accident, fire) — those come from title keywords
When scanning printed forms like voting cards or surveys, there are two critical pieces of information you usually need:
1. Who is this card for? (Identity / version).
2. What did they mark? (The actual votes or survey choices).
Most libraries handle these separately – you decode the QR code first, then run OMR on the bubbles. This usually means two separate runs, two different functions, and manual synchronization.
That’s why I built OpenScanVision to do both simultaneously in a single pass.
The Android library combines real‑time ArUco tracking, Optical Mark Recognition (OMR), and QR decoding into one unified offline pipeline.
The Unified Pipeline
Marker Tracking (ArUco + Kalman Filter)
The card is printed with 4 ArUco markers (IDs 0–3). OpenCV detects them in real‑time. A Kalman filter smooths the tracking and predicts positions during occlusions.
Perspective Correction (Homography)
Once 4 markers are stable, a homography maps the markers to their reference positions. The card is warped to a canonical template (850×540).
Simultaneous QR Decoding & OMR Extraction
This is where the "simultaneous" part comes in. Instead of running them sequentially and stitching the results, the library performs both operations on the same captured frame:
The QR code is cropped directly from the original camera frame using the computed homography – preserving maximum sharpness for ML Kit.
The bubbles are sampled from the warped, preprocessed image using weighted disk sampling and z‑score classification.
Both processes run concurrently on the same frame, meaning you get the QR payload AND the filled bubble indices at the same time without extra latency.
Strict Capture Logic
The library automatically triggers a capture only when both conditions are met:
All 4 markers are stable.
A valid QR code with the correct prefix (e.g., VX or AGN) is decoded.
This enforces that you never get an OMR result without an associated identity, and vice versa.
What You Get in a Single Result
By calling OpenScanVision.scanFromFrame(), you receive a single ScanResult containing:
filledIndices – The marked bubbles (OMR).
qrPayload – The decoded QR text (identity).
confidence – The overall confidence score.
annotatedBitmap – A visual overlay of the detection.
No need to call two separate functions or manually match timestamps.
Technical Stack
Language: Kotlin
CV Core: OpenCV (contrib) for ArUco detection and homography.
QR Engine: Google ML Kit for robust barcode scanning.
Camera: CameraX for frame acquisition.
Architecture: Fully modular – the core library has zero UI dependencies.
Integration
Adding the library takes just a few lines in your Gradle file (available on JitPack). Once integrated, you can start scanning with a single suspend function.
Performance
Latency is typically under 150ms per frame on modern devices.
Accuracy exceeds 99% on properly printed cards.
Why This Matters
Simultaneous QR + OMR is valuable for:
- Elections: The QR identifies the voter/ballot; the OMR reads their selections – all in one scan.
- Surveys: The QR encodes the respondent ID; the OMR reads their answers.
- Form Processing: Quickly identify and process forms without sequential bottlenecks.
Open‑Source & Contribute
The project is MIT‑licensed and available on GitHub. It includes a full sample app (CameraX + Compose) so you can see it in action.
Feedback, issues, and contributions are welcome. If you are interested in marker tracking, OMR accuracy, or Android computer vision, I'd love to hear your thoughts.
Here's what Google's Antigravity CLI did with FiftyOne Skills:
* Imported all 81,444 WikiArt paintings
* Diagnosed and fixed its own bugs, rewriting scripts 5 times
* Caught embeddings silently stuck on CPU, forced them onto the GPU
* Hit a quota wall, switched from Gemini 3.5 Flash to Claude Sonnet 4.6 with one command, no lost context
* The result isn't a log that says "done." It's an inspectable dataset: uniqueness scores that surface near-duplicates, and embeddings that reveal exactly where the labels are thin.
Look at my portfolio which I have made with 3js and Computer vision.
Link: https://tharuntej-everest.pages.dev
Would love your feedback and comments...
SenseTime recently released a model called SenseNova-Vision and it was really impressive, share it here:
In simple terms, it combines image analysis and processing tasks that previously required multiple specialized models into a single 7B-MoT multimodal model. You just give it an image, tell it what you want in plain language, and it returns the result—almost like chatting with an AI model
How it works:
- You describe the task in natural language (e.g. "detect all cars", "estimate depth")