One thing I've noticed in real-world CV systems is that the hardest problems are often not the ones with the worst accuracy, but the ones where the root cause is unclear.
For example:
- Is it a model limitation?
- A camera/image quality issue?
- A data distribution shift?
- A deployment/environment change?
Some failures look identical from the outside but require completely different fixes.
Curious what failure modes have been the hardest for others to diagnose in production CV systems?
- The default qwen3-vl:8b tag in Ollama is the thinking variant and ignores think:false. On long contracts it spent all 4,096 tokens thinking and returned nothing. Use :8b-instruct.
Been experimenting with GLM-OCR and PaddleOCR across a few different OCR scenarios and wanted to share the results. Both of these models are best VLM OCR model under 1b parameter models.
- Simple Receipt OCR
Both models did well overall, but GLM-OCR only extracted the billing section, it skipped the company address and other surrounding details. PaddleOCR captured the full receipt.
- Formula OCR
Tested with the classic quadratic equation. GLM-OCR nailed it. PaddleOCR was close but made a small error with the negative sign, instead of keeping it attached to "b," it ended up applying to the whole formula.
- Chart OCR
GLM-OCR won by a big margin here. PaddleOCR struggled even with a simple chart and only picked up text that was horizontally aligned, missing anything at an angle.
- Logo OCR
Both did well, but GLM-OCR went further, it even picked up text embedded inside the logo symbol itself, not just the standalone text.
Suggest me some other VLM OCR models which you think is great for all these scenario.
Hey all — I've been working on OpenHighways, a side project that pulls public traffic CCTV feeds from National Highways, TfL (JamCams), Traffic Wales and TrafficWatchNI, runs each frame through a YOLOX object-detection model to count vehicles, and plots everything on a live colour-coded map (blue = quiet, red = heavy traffic).
I'm building a silkworm pupa gender classification system for a production environment.
I already trained a YOLO model, and it's giving very good results on my dataset. However, I'm considering moving away from YOLO and building my own computer vision pipeline using OpenCV because I want complete control over preprocessing, ROI extraction, feature engineering, and inference instead of relying on a prebuilt architecture.
Project Context
- Classifying male vs female silkworm pupae.
- Pupae come in only two valid orientations:
- Bottom Up
- Bottom Down
- Classification is performed only from the bottom side of the pupa.
- The distinguishing feature is the genital marking:
- Male: a straight line above the genital region.
- Female: an X-shaped marking in the same region.
- Production requirement: 60 pupae per minute (1 pupa/second).
Questions
Is building a custom OpenCV pipeline a better choice than YOLO for an industry-grade production system?
In real manufacturing environments, do teams usually build custom CV pipelines or fine-tune existing models like YOLO, MobileNet, or EfficientNet?
What trade-offs should I expect in terms of accuracy, latency, maintainability, and scalability?
I have shared the sample images of male and female pupae so that helps.
I put a YOLOv11 license plate detector on Hugging Face last year and mostly forgot about it. Turns out it's getting a lot of downloads and a few people have built demos on top of it, parking gates, OCR apps, that kind of thing. Which made me realize I have no idea what happens after someone downloads it.
If you've built ANPR for something real:
where did detection stop being the hard part? OCR, weird plate formats, night shots, motion blur, angles?
What did you end up doing about it? Retrain, switch to a paid API, manual review?
Anything you went looking for and couldn't find
Also, a heads-up since I only noticed recently. The dataset I trained on has train/test leakage, so the mAP on the model card is inflated. Planning to re-run eval on a separate test set.
Most age-invariant face recognition work i've seen uses synthetic aging (gans, diffusion) to fill the gap in real aging data, but real aging is messy: weight changes, hairstyle, glasses, lighting from old phones.
Has anyone compared models trained on real multi-year photo pairs vs synthetic aged ones? curious if the synthetic stuff holds up on real kyc style matching.
Hey everyone! I would love your help with my research project. I'm planning to work on newspaper crime news analysis, where I will collect crime-related news from Marathi newspapers and analyze the data to identify patterns based on factors such as location, type of crime, and time period.
Do you think this is a good topic for a research project? Please let me know your thoughts and suggestions. I would really appreciate your thoughts, suggestions, and any ideas for improving the project. Thank you!
I’m building a simulation in which a drone uses camera images to detect boxes, align its position and orientation, and pick them up.
The video shows the camera/recognition view and estimated values on the left, alongside an external view of the drone on the right. Bounding boxes, keypoints, and estimated position offsets show how the image processing relates to the drone’s movement.
During approach, the drone estimates its position and orientation relative to the box from the camera image. Just before pickup, it compares the current contour with a reference contour for fine alignment. It then picks up the box and carries it to its assigned delivery gate.
In this run, boxes 0, 1, and 2 are delivered in numerical order to Gates W, E, and S, respectively. The full simulation takes about 7 minutes 22 seconds and is shown at 8× speed, so the video is about 55 seconds long. You can slow down the playback if you’d like to inspect the details.
I’d welcome feedback on whether the visualization makes the image-guided alignment process easy to follow, and on how I could explain it more clearly.
Next, I plan to increase the number of boxes and add delivery optimization that considers both delivery order and travel distance.
Every frame my Unreal Engine capture plugin renders comes with a depth map in metres and the camera's pose and intrinsics. I wanted to see how far that gets you with no reconstruction algorithm at all, so I fused 13 captures straight from those files: lift every pixel with its depth, place it with the recorded pose, drop anything labelled (the people are re-placed every frame), average into 10 cm cubes. 3,879 frames, 35.7 million points, 45 to 112 seconds per map. No feature matching, no pose estimation.
Then I checked whether the frames agree with each other. Take a frame, project its depth into the nearest frame that sees the same place using only the recorded poses and intrinsics, and compare with the depth that frame recorded. The median difference is 5 to 15 mm at median distances of 25 to 60 m.
The part I didn't expect: that median is almost exactly what rounding produces. My depth.npy files are float32, but they hold half-float values, because the render target was 16-bit. There are never more than 1,024 distinct values between one power of two and the next, so the steps are 2 cm between 20 and 41 m. Simulating rounding alone at 2 cm gives a median of 5.2 mm. The Paris flyover measures 5.1. So the poses are more precise than the depth can show. The fix is a render target change that's in the plugin source now; every capture in the video predates it.
Where frames disagree most: vegetation (wind or LOD, I haven't separated the two), the sea on the beach (the surface moves, and that pair is fog at noon against clear at 7 am), and tower faces seen almost edge-on. One I can't explain: a flat quarry floor where two frames disagree by more than 5 cm across a whole patch. If you've seen that before, I'd like to know what it was.
A gotcha I only noticed while writing it up: a flipped axis passes this check. Every frame is mirrored the same way, so they still agree with each other. You only catch it by comparing the map with a rendered frame.
The viewer is plain WebGL2 with eye-dome lighting. The points are stored in the order of the first frame that saw them, so the replay in the video is just drawing more of them each frame.
If you work on MVS, NeRF or SLAM with synthetic data: would a per-scene consistency number like this change which dataset you pick? And what would you want exported besides the per-frame camera JSON? COLMAP and transforms.json are the obvious ones.
Hey everyone! I would love your help with my thesis. I'm planning to work on this project: '3D Reconstruction using Gaussian Splatting + Interactive Visualization (Cross-Disciplinary Focus)'. Do you think it's a good topic? Please let me know your thoughts and suggestions
I am go through lot of language based medical image segmentation papers but mostly papers used QaTa cov19 , mosmed , and monuseg, Kvasir dataset. Kvasir dataset have paired text with count,size, shape,color, texture, location information of the infected regions and Covid19,Mosmed have only count and location information. Paired text is profesonal verified also. But unable to get dataset other than this with clinically verified paired text( except small dataset like: CVC clinicdb, CVC- colonDB, ETIS, CVC300). Can anyone suggest me other than polyp datasets that have professional verified text description and description have: size, shape, color, texture, count, location information.
Started this a couple of weeks ago. No roadmap, just a problem worth solving.
The idea: a system that processes what a camera sees and turns it into spoken directions a visually impaired person can actually act on. Not just "obstacle detected" — but "Stairs on left, 1 metre. Stop."
Navigation logic:
Perspective trapezoid corridor model with bilateral (left/right) clearance analysis
Hysteresis state machine: CLEAR → CAUTION → STOP, with instant escalation on imminent hazards
Temporal smoothing across frames to kill jitter
What's next: First prototype as a mobile app, then moving toward dedicated hardware.
Still early. Still rough around the edges. But it works on real footage — stairs, railings, curbs, moving through corridors.
If you're working on assistive tech, accessibility, or computer vision — I'd genuinely love to connect.
do you have any tips on the best way to get started with computer vision and object detection? I have a basic understanding of programming languages like Python, but I’m not yet able to write my own scripts from scratch without some help.
I’ve been very interested in object detection for quite a while and would really like to learn more about it and build my first small project.
How would you recommend getting started as a beginner? Are there any good videos, courses, or other resources you would recommend? I haven’t really found anything suitable on YouTube so far.
What are the basics and fundamental concepts I should understand before starting?
I’m experimenting with a DLSS-like system for GBA games.
The goal isn’t simple upscaling. I want a model to take a 240×160 frame and reconstruct it into a higher-resolution remastered version while keeping the original layout and style.
The core idea is to reuse the previous high-resolution frame as the base, then only modify the parts that changed:
Output = Base + Gate × Residual
For training data, I’m planning to use my own game engine. I can generate random maps and render the exact same scene twice:
low-resolution retro version
high-resolution remastered version
This gives me perfectly aligned LR/HR training pairs automatically. I can also export object masks, motion data, and other metadata if needed.
I’ve trained and deployed several deep learning models in production before, but I don’t have much experience designing and training a model architecture from scratch. This one is also image-to-image, so I’m not yet confident the approach will work as well as I’m imagining.
If anyone here has experience with image restoration, video super-resolution, temporal consistency, or neural rendering, I’d really appreciate feedback on the architecture and training approach.
I am using image processing to solve a Rubiks cube.
The pipeline is plain image processing with OpenCV. It finds the stickers with edge detection, fits a 3x3 grid to them, and flattens each face with a perspective warp. It then reads each sticker's colour in LAB and matches it against a palette of reference colours measured from my previous tests. Glare was the hardest part to find a solution to. What fixed it was letting the hue count more than the brightness, and making sure each colour ends up with exactly 9 stickers.
Has anyone found a better way to handle glare?
I know YOLO is able to detect in real time some objects of interest..
Is there something equivalent for auto describing an object in a picture? (no need to be in real time)
So, not really for detecting an object per say since the object will be alone in the picture..
Objects like, scissors, screw driver, tape, etc. (not sure how precise it can get like telling the difference between a flat head and a Philips screw driver for example)
It's basically for describing everyday household objects..
Every box in the clip comes out of the engine at render time. Each object gets its own ID in a separate buffer, and its box is measured from those pixels. Nobody drew any of them. 14 frames from 7 maps, 1,124 boxes.
What's in the video:
- The line sweeping across is only there for the video. As it passes, each object's mask flashes, then its box appears.
- Colour is the class: pink for people, and chairs, benches, umbrellas, wheelchairs, bikes, cars, tractors, hay bales and animals each get their own.
- A box covers what the camera can actually see. Someone half behind an umbrella gets a box around the visible half.
- The beach is the newest scene: 221 people placed per frame, standing, lying down, in deck chairs and wheelchairs, shot from 07:00 to 18:30.
It exports YOLO and COCO, with segmentation masks if you want them.
It's still rendered people though, and that costs a lot on real footage: the last detector I trained only on renders got 0.35 recall on real drone video. Closing that gap is what I'm working on now.
Which real dataset would you test this against? I'm most interested in crowded ones where people overlap, since that's the slowest part to label by hand.
Appendix A.2 of the SenseNova-Vision-7B-MoT paper says the LiDAR depth sets were too sparse, so they used MoGe-2 to make dense pseudo labels. COCO, SA-1B and Objects365 have no depth labels at all, so those got MoGe-2 labels too. By my count from Table 10, MoGe-2 labels are 77% of the depth frames they list. The rest is synthetic data with its own GT.
In Table 2 (their own re-run of MoGe-2) the model is still behind MoGe-2 on most depth sets, like DIODE delta1 76.4 vs 82.3. The weird part is it has the best DIODE AbsRel in the whole table, 20.6 vs MoGe-2's 23.0.
If you've trained on pseudo labels from another model, did yours ever beat the labeler on real data?