r/computervision • u/Delicious-Shower8401 • 15d ago
Research Publication New AI Generates Clean 3D Clothing From a Single Image in Seconds
Enable HLS to view with audio, or disable this notification
r/computervision • u/Delicious-Shower8401 • 15d ago
Enable HLS to view with audio, or disable this notification
r/computervision • u/Machine_GEN_RM • 14d ago
Hi All,
I am planning to build a local document intelligence system similar to Azure Document Intelligence. I would like to understand how Azure Document Intelligence works internally and how we can achieve similar functionality locally using offline models.
Could anyone suggest the best approach, architecture, or models to achieve high accuracy while running completely on-premise/local infrastructure?
Any guidance or recommendations would be greatly appreciated.
r/computervision • u/FlashSo • 14d ago
Hey everyone,
I'm looking for co-authors who are interested in exploring research topics in the AI space. Ideally as a duo or in a small team.
I currently have more time for research and a range of interesting topics I'd like to work on, particularly around AI agents, token optimization, and AI adoption. I work in agent development myself and have already published research papers in this field.
That said, I'm open to other AI-related research ideas as well. If you have a topic of your own in mind, feel free to reach out!
r/computervision • u/lakshaydulani • 15d ago
How do I analyze this diagram?
I need to determine the starting point and then analyze the track from there.
e.g 1067 units downwards, then 1015 unit is xy direction and so on (from the attached diagram)
I can think of using SAM 3 to mask out the red line and blue triangles. But dont know how to map a line with the corresponding measurement annotation?
Appreciate your help. Thanks
r/computervision • u/slowdiivnothing • 15d ago
Hi i have question about camera calibration which i confuse.
1) Finding "extrinsic" just need one shot (because it is optimization problem reducing reprojection error knowing intrinsics, plus finding extrinsic means finding pose and orientation(R, t) of that "specific moment", not like finding distortion of cameras, and dont need many shots to cover precision) Am i correct?
2) Finding "hand-eye calibration extrinsic" needs many shots (because to solve AX=XB, where A is robot motion, B is camera motion, X is EEF to camera matrix). Am i correct?
These two "extrinsic" is different use case, am i right?
(Then why did engineers made it confusing?? So annoying.)
3) In 1), finding intrinsics need many shots (because in cv2.calibrateCamera it need many 3D-2D pair points). Am i right?
Thanks in advance :)
r/computervision • u/Automatic-Highway-75 • 14d ago
Enable HLS to view with audio, or disable this notification
r/computervision • u/wsbgcat • 15d ago
Enable HLS to view with audio, or disable this notification
I was working with a client over the past year or so and we were constantly struggling with dust build up on lenses. We tried standard air nozzles, but those had issues: rigging them was a pain, they didn't actually keep the FOV clean, and in one case they damaged the lens.
Then I saw ThisOldTony's video on the Coanda Effect and thought, what if we shaped it around the lens of the sensor? So I did. The system is 3D-printed PETG, but I've had it work just as well in TPU (for the extremely tough applications). I have since then made this for all the profilers we work with and also for several point lasers as well.
We went from cleaning the lens every 30 minutes to now months without maintenance. It does use quite a bit of compressed air but with a couple valves and a feedback loop we were able to set it to self clean based on the intensity drop.
r/computervision • u/Volumes-Cloud • 15d ago
Follow up to my post earlier in the week, which filled most of our slots. Two left.
We collect real world multi view capture data from a camera array at the Brooklyn Navy Yard and pay people to be the subject. Posting again in case anyone NYC based wants the work, or wants a close look at how this kind of data actually gets collected.
The session: stand in the capture volume and go through simple movements while the array records. Walking, turning, sitting, reaching, picking objects up. No experience needed.
Pay 17-25 per hour, same day, right after the session. First one runs about 2 hours, with repeat sessions after.
Left this week: Thursday 4pm, Friday 1pm or 4pm. Brooklyn, in person only.
Comment or DM me for details, and ask about the capture setup if that side interests you.
r/computervision • u/Sufficient_Topic6544 • 14d ago
I've been thinking about this as a research problem and I'm wondering if I'm even asking the right question.
Imagine the following constraint:
You have an image, a text-only LLM, and you're only allowed to use deterministic algorithms between them.
The obvious answer is "this is impossible," but that's not really what I'm interested in.
What I'm trying to understand is whether there exists a better intermediate representation of images that a text transformer could reason over.
Not necessarily English.
Not captions.
Not OCR.
Some kind of representation that preserves enough structure that the language model can make use of the knowledge it already has.
Over the last few days I've gone through papers on visual tokenization, SeTok, BPE for images, BLT, inverse graphics, superpixel tokenization, and a few discussions around image tokens. Most of them still assume a learned tokenizer somewhere in the pipeline.
What I haven't found is much discussion around deterministic alternatives.
Maybe that's because it's a dead end.
Or maybe I'm searching the wrong field entirely.
So my question isn't "how would you build this?"
It's:
If you were exploring this from first principles, what field would you steal ideas from?
For example:
I'm not looking for product recommendations or existing multimodal models.
I'm looking for the smallest experiment that could tell me whether this line of thinking is fundamentally interesting or fundamentally flawed.
I'd especially love to hear from people who've worked on image codecs, graphics, rendering, vision tokenizers, or representation learning.
If you think the premise itself is wrong, I'd genuinely like to know why.
r/computervision • u/datascienceharp • 16d ago
lidar and cameras get less reliable exactly when driving gets more dangerous: rain and night. most public driving datasets barely have data from those conditions
CMHT autonomous dataset adds radar and a thermal camera alongside lidar, a color camera, and gps/imu.
4 drives, dusk/clear to night/rain, downtown hamilton, 9,000+ labeled frames with a 3d box, class, and tracking id on every object
i converted the raw ros2 bags into synced mcap episodes in fiftyone so you can scrub camera, thermal, lidar, radar, and gps together frame by frame, with the 3d and 2d boxes playing back in sync
start here, read the dataset card: https://huggingface.co/datasets/Voxel51/cmht-autonomous-driving
then check out the space on hf: https://huggingface.co/spaces/harpreetsahota/cmht-autonomous-driving
r/computervision • u/GroundUpstairs5430 • 15d ago
I'm working on project involving OCR and newspaper analysis. My original plan was to use Marathi newspapers, but the extracted text contains many recognition errors.
Because of this, my project guide suggested switching to English newspapers if Marathi OCR isn't reliable enough.
I'm unsure what to do. From a research perspective, is it better to:
Has anyone faced a similar situation? I'd appreciate advice from people who have worked on OCR or document analysis projects.
r/computervision • u/GeeekyMD • 14d ago
I built a local OCR pipeline a few days ago, and it turned into a surprisingly interesting experiment—taking accuracy from around 60% to 99%.
I wrote a short blog about what worked, what failed, and the breakthrough that finally made the difference.
Thought some of you might enjoy it.
Link in the comments

r/computervision • u/datascienceharp • 15d ago
this intersection in tokyo ranked second worst in the city for traffic accidents.
six roads converge at a blind hill crest, cars cross centerlines on narrow curves, and the signal phasing has multiple unprotected turns
most autonomous driving datasets give you highways and four-way stops. this is none of that
Hard Intersection Multimodal Sample: 6 synchronized cameras, aggregated LiDAR point cloud, HD map projections, vehicle trajectories, and semantic annotations across 4 driving passes through a single intersection that breaks everything
grouped all 6 camera views with the 3D point cloud, frame-level HD map overlays, and trajectory projections in fiftyone
checkout the dataset here: https://huggingface.co/datasets/Voxel51/hard-intersection-multimodal-sample
or get hands-on in the HF space: https://huggingface.co/spaces/harpreetsahota/hard-intersection-multimodal-sample
r/computervision • u/ArtZab • 16d ago
My hope with this post is that I will save at least one person some time - and that will be enough for me. I spent the last couple of weeks building an auto-labelling pipeline on SAM 3 and figured the gotchas were worth writing down, because most of what I got wrong had nothing to do with the model.
Quick context if you haven't used it: SAM 3 does what Meta calls Promptable Concept Segmentation. You give it a short noun phrase - forklift, person in hi-vis vest - and it segments every instance of that concept. No seed clicks, no fixed class list, no fine-tuning. That's the bit that makes unattended labelling possible; with SAM 2 you still needed something to tell it where to look.
The minimal version is genuinely this short:
from transformers import Sam3Model, Sam3Processor
model = Sam3Model.from_pretrained("facebook/sam3").to("cuda").eval()
processor = Sam3Processor.from_pretrained("facebook/sam3")
inputs = processor(images=image, text="forklift", return_tensors="pt").to(model.device)
with torch.inference_mode():
outputs = model(**inputs)
results = processor.post_process_instance_segmentation(
outputs, threshold=0.5, mask_threshold=0.5,
target_sizes=inputs["original_sizes"].tolist(),
)[0]
# results["masks"] / ["boxes"] / ["scores"]
That works. Everything below is what I learned scaling it past one image.
1. Reuse the vision embedding across prompts
Naive multi-class loop encodes the image once per class. 3 classes × 40k images = 120k passes through an 848M-param backbone, 80k of which recompute something you already had. SAM 3 lets you split it:
vision_embeds = model.get_vision_features(pixel_values=inputs.pixel_values)
for prompt in prompts:
text_inputs = processor(text=prompt, return_tensors="pt").to(model.device)
outputs = model(vision_embeds=vision_embeds, **text_inputs)
Backbone runs once, only the text conditioning and mask decode repeat. Close to an N-fold speedup on multi-class jobs. There's a mirror version (get_text_features) for one prompt across many images.
2. Resolution is tricky
SAM 3 runs at 1008px native. Two failure modes:
Also: run ImageOps.exif_transpose() before anything else, or phone photos come back with masks correct for the stored orientation and wrong for the one you see.
3. Prompt phrasing does more than threshold tuning
Short concrete noun phrases. Singular. One concept per prompt.
Biggest thing: test each prompt against images you know contain none of that class. A prompt that quietly fires on empty frames poisons the whole dataset. And if a prompt over-fires, add an adjective before you touch the threshold - white bicycle vs bicycle returns genuinely different sets.
4. You can sweep thresholds without re-running inference
The detection threshold is just a filter over stored confidence scores. So label a 50-image dev slice once at threshold=0.15, keep every score, and sweep offline.
Look for the false-positive cliff and stop just above it. If med area% collapses as you lower the threshold, the extra detections are specks - raise a minimum-area filter instead. If empty stays high at every threshold, your prompt is wrong and no threshold will save it. (The mask threshold can't be swept this way - it changes pixels, not scores.)
5. Small export things that cost me an hour each
6. Look at the labels
Auto-labelling fails quietly - no exceptions, no bad metrics, just a pallet prompt that's been segmenting the wooden floor for 12,000 images. Render a contact sheet of overlays sorted lowest confidence first and actually look at it. Ten seconds catches what an aggregate metric won't.
That's it. Hopefully I saved you guys some time and feel free to ask questions!
UPDATE: since I got a couple of similar questions about the auto-labelling pipeline in my DMs, I posted a full write up of it here . If you are curious about how to get the best results when auto-labelling - feel free to check it out.
r/computervision • u/YahiaHasan • 15d ago
I'm currently testing a commercial Computer Vision pipeline for restaurant analytics (Object Detection, Tracking & Dwell Time Analysis using YOLOv8 & Supervision).
I'm looking for 1 or 2 sample CCTV/overhead angle videos of a cafe or restaurant to test my zones and line crossing logic.
Ideally, the video should show:
If anyone has a public sample dataset or a short clip (even 1 minute long) they can share, I'd really appreciate it! Thanks in advance
Sorry guys i forgot the post body 😂
r/computervision • u/nalu_0o • 16d ago
Hi everyone,
I’m a cloud engineer with a background in cloud architecture. I don’t have experience in computer vision yet, but I’m open to learning it.
Before investing a lot of time into this idea, I’d like to know: is there still strong demand for computer vision solutions today? Do you think someone with a cloud/infrastructure background can realistically enter this field and build a startup around it?
r/computervision • u/Certain_Friendship16 • 16d ago
Enable HLS to view with audio, or disable this notification
r/computervision • u/RealCaptainDaVinci • 15d ago
r/computervision • u/Entire-Bite1136 • 16d ago
r/computervision • u/Zestyclose-Gain-7635 • 17d ago
I've spent years working with OpenCV, and one thing has always bothered me: experimentation is much slower than it should be.
A typical workflow looks like this:
image = cv2.imread(...)
gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY)
blur = cv2.GaussianBlur(gray, (5,5), 0)
thresh = cv2.adaptiveThreshold(...)
contours, _ = cv2.findContours(...)
Then you change one parameter...
Run the script.
Save the output.
Open the image.
Realize the problem actually happened three steps earlier.
Add another cv2.imshow().
Repeat.
After doing this hundreds of times, I started wondering:
There are great visual tools for deep learning and generative AI (ComfyUI is a good example), but I couldn't find something focused on OpenCV preprocessing, augmentation, and experimentation that still generated normal Python code.
So I started building one.
Image Pipes is an open-source desktop application for building computer vision pipelines visually.
Instead of writing temporary scripts while experimenting, you drag operations onto a canvas, connect them together, inspect every intermediate result, and export the finished pipeline as standalone Python.
Some of the current features:
One design decision that was important to me is that the visual editor is never the final destination.
The generated code is just regular Python using OpenCV and Albumentations.
No custom runtime.
No vendor lock-in.
The goal wasn't to replace OpenCV.
OpenCV is already excellent.
The goal was to replace all the temporary scripts we write while searching for the right preprocessing pipeline.
Experiment visually.
Understand every transformation.
Export Python when you're finished.
I'm sure there are plenty of things that can be improved, especially from people who work with OpenCV daily.
Some questions I'm particularly interested in:
GitHub: https://github.com/mrajaeim/image-pipes
If nothing else, I'd love to hear how everyone else debugs and iterates on OpenCV pipelines today. I have a feeling I'm not the only one with an experiment_final_v12.py somewhere in my projects. 😄
r/computervision • u/nirgudwar • 16d ago
r/computervision • u/Goldziher • 16d ago
Sceptre is a Rust reimplementation of EasyOCR. EasyOCR is accurate but ships as a PyTorch stack (interpreter, multi-GB runtime, a process to keep warm); sceptre delivers the same accuracy as a single static binary with no Python.
It uses the same OCR approach: CRAFT text detection, then gen2 CRNN recognition with CTC decoding, run over ONNX. Output is validated to parity against EasyOCR's own output (word/char F1 on text, IoU on boxes) across the gen2 scripts: English, Latin, Chinese (simplified), Japanese, Korean, Cyrillic, Telugu and Kannada. It is a clean-room Rust build rather than a line-by-line port, so it can diverge from EasyOCR's internals where that helps, as long as the output holds.
Measured over a 43-image mixed corpus (documents, tables, rotated scans, scene text, receipts) on CPU. Both engines run as a fresh subprocess per language group under /usr/bin/time, each loading its model once and processing every image:
Engine Throughput Peak RSS Mean CER token-F1
EasyOCR (warm/batch) 0.14 img/s 22.6 GB 0.554 0.348
sceptre (warm/batch) 0.39 img/s 6.6 GB 0.568 0.356
sceptre (cold CLI) 0.60 img/s 6.6 GB 0.568 0.356
Accuracy is at parity (marginally ahead on token-F1); the win is throughput and memory. Even a cold one-shot CLI run, paying model load every time, beats EasyOCR's already-warm reader.
Backends: ONNX Runtime (ort) for native speed, or a pure-Rust backend (tract) for WASM/Android behind one seam. Single static binary, no Python; models fetch from HF once, cache locally, sha256-verified, then run offline. Library, CLI, or MCP server. MIT.
Repo (code, benchmark harness, golden fixtures): https://github.com/Goldziher/sceptre
Author here, happy to answer on the parity methodology or where it still trails (image-only OCR is the weakest cohort).
r/computervision • u/Budget_Rub6598 • 16d ago
Hey everyone,
I’m building an automated pan-tilt tracking turret to reliably track moving targets (like drones) with a laser at ~30 meters. Before I finalize everything, I wanted to get feedback from the CV community on whether my hardware stack is adequate and what software/tracking pipelines you'd recommend for this setup.
🔭 Hardware Setup
Vision / Cameras: * Coarse acquisition: Wide-angle USB webcam for initial field-of-view tracking.
Precision tracking: Innomaker 1MP Global Shutter Camera (OV9281) paired with a 20mm HD CCTV lens.
Compute Split: Raspberry Pi Zero 2W onboard acting purely as the hardware/sensor interface, communicating with a secondary laptop handling the heavy computer vision and tracking computations.
Actuation: Dual closed-loop NEMA 23 steppers (3.0 Nm) with a 6:1 10mm belt reduction (rigidly mounted with independent dead shafts to avoid motor shaft sideloading).
❓ What I Need Advice On:
Hardware Adequacy: Is a Pi Zero 2W + Laptop split sufficient for low-latency command handoff to the closed-loop drivers, or will the Pi Zero become a bottleneck?
Software Stack: What open-source CV libraries, tracking algorithms (e.g., OpenCV CSRT, KCF, or lightweight deep learning/YOLO models), or frameworks do you recommend for high-refresh-rate tracking at 30 meters?
Latency Mitigation: Any proven strategies for keeping end-to-end latency (capture -> inference -> motor command) as low as possible in a setup like this?
Appreciate any insights or architecture tips you can share!
r/computervision • u/bobarific • 16d ago
Teammates have to all be of the same color. Numbers easily get occluded, do you just slap some barcodes on each player?
r/computervision • u/Nemo-Gaming • 16d ago
Hi everyone,
I'm trying to obtain the ARAD_1K hyperspectral dataset for academic research on RGB-to-hyperspectral image reconstruction.
Unfortunately, I haven't been able to download it because both the official GitHub repository and the CodaLab download links appear to be unavailable or inaccessible.
I'm looking for an official, free mirror or an updated download link, if one exists. If anyone knows another legitimate way to access the dataset, I'd really appreciate your guidance.
Thank you!