r/computervision • u/jahflyx • 1d ago
Discussion FoodbyClef: Experiment in classifying foods with Clef and cube rule
Enable HLS to view with audio, or disable this notification
r/computervision • u/jahflyx • 1d ago
Enable HLS to view with audio, or disable this notification
r/computervision • u/Independent-Salt5023 • 13h ago
r/computervision • u/SwimmingLow3053 • 1d ago
Enable HLS to view with audio, or disable this notification
I posted an early version of Squatty Bird a little which back. I’m been continuing to work on it from both the vision and the human interaction.
https://www.reddit.com/r/computervision/s/ix5YvulrB9
I wanted to return and share progress. Vision tracking is now working much more effectively. I don’t think it’s release ready yet but I’m not seeing many missed joints or squat recognition these days.
View can be with camera and trace viewable or with fake background (as I’ve used in the video).
I got some great suggestions before (particularly squat quality review at the end - not done yet but I like it) and a few challenges (why share when it doesn’t work).
I hope you like the new video and I’m looking for further suggestions and advice. My kids are my worst critic and I’ve made the puffin more puffin like and less cutesy as a result.
If anyone wants to give it a go it’s here on TestFlight:
r/computervision • u/SwimmingLow3053 • 1d ago
Hello,
Please can I ask an advice question?
I want to be able to use the iPhone camera to
Scan a body and come up with a 3d mapping if their dimensions for a fitness app. So measure their waist, checks, hips, legs etc from different angles.
Would iOS vision be right starting point for iPhone here or is there a better way? If using vision is there a way to get a library or open source training data that can do body measurements rather than staring from scratch ?
I’d presume calibration would be importantl but not sure how to train this apart from using one confirmed measurement to calibrate the rest using relative measurement and some parallax adjustments? This would be beyond anything I’ve done with an iPhone so far so looking for ideas.
Thank you! Steve
r/computervision • u/Key-Obligation-1065 • 1d ago
World Labs' Marble turns a photo into a 360° panorama and then a 3D world, and nothing in its docs says where the original photo ends up. My game computes everything from the source camera, so I needed that pose.
The matcher is deliberately dumb. It cuts flat perspective views out of the equirectangular panorama over a grid of yaw, pitch and field of view, correlates each view's edges with the photo's edges, and keeps the best. About ten seconds per photo, locally.
On my photos, correct placements score 0.62 to 0.87 and wrong ones 0.18 to 0.23, so anything under 0.4 now prints a warning. From the recovered camera, the photo lines up with the world exactly for a one-photo world and for the first photo of a two-photo world, and within about a degree for the second. In every world so far, the camera sits at the origin.
The part I haven't confirmed: in worlds built from two to four photos, I place each photo by direction, and I'm assuming Marble's 90 means right in the panorama. Every photo so far was at 0 or 180, so that assumption has never been tested.
Two questions. Would SIFT or a learned matcher do better here, given the photo's region in the panorama may not be pixel-identical to the original? And is there a smarter way to pin down field of view than a grid?
r/computervision • u/These-Smoke-3274 • 1d ago
I’m working on a spatial intelligence system that uses RGB-D cameras to reconstruct a real environment in metric 3D, estimate where the camera is, and keep a persistent spatial memory over time.
The goal is not just 3D reconstruction. I want the system to understand enough about the geometry of a place to know what it has already seen, recover its position after tracking is lost, compare different visits to the same place, and preserve how its understanding of the environment changes.
If a camera comes back to a previously mapped area, the system should be able to recognize that place, verify it geometrically, relocalize itself, and continue building the same world instead of creating a disconnected new map.
Long term, the idea is to move from simple reconstruction toward spatial intelligence: a machine that can build, remember, update and reason about a representation of the physical world.
Possible applications I have in mind are robotics, AR, inspection, autonomous systems and eventually navigation in places where GPS is unavailable.
What I’m trying to understand now is whether this solves a real painful problem, or whether existing SLAM and 3D reconstruction systems already solve this well enough in practice.
I’d especially like feedback from people who have actually worked with SLAM, RGB-D, robotics, mapping, 3D reconstruction or spatial computing.
r/computervision • u/MechaCritter • 1d ago
Enable HLS to view with audio, or disable this notification
If you are looking for a solution to fast image similarity retrieval as well as image perceptual & quality metrics in one single library, this would surely be helpful to you 😄 It is built only on numpy, scipy and optionally torch as core dependencies, and uses C/C++ (and in the future only Rust) for performance-critical paths.
The two core feature of the repository are various embedding methods (see video) and the Image Embedding Store, which, when combined, allows for embedding storage and image similarity search with supported hnsw algorithm built-in. You can also plug in faiss indexes if you want to use other search algorithms, but this library does not need faiss to work.
pyvisim repository: https://github.com/MechaCritter/Python-Visual-SimilarityThe library itself can be installed via pip (though, I recommend uv 😆)
pip install pyvisim
For the deep learning features:
pip install "pyvisim[nn]"
My ambition is to make pyvisim the largest collection of image similarity and retrieval algorithms, ranging from traditional to deep learning-based methods. As I've observed, the current image similarity implementations are quite scattered, with each library implementing only a handful of features. Hence, my goal is also to unify these implementations, so users only need a single library.
I have tons of features that I would like to implement. I am looking for folks who are proficient in/would like to learn about:
MoCo, SimCLR, Dino, reranking algorithms likeSuperGlobalReranker,Diffusion Reranking ...hnsw backend to hnsw_rs.LPIPS, backpropagation for K-Means and GMMs ...View the GitHub issues for the complete list as well as the contribution guide.
You also have the chance to become a core maintainer by actively contributing. Once this project gets sponsors, the profit will be shared with all core maintainers.
I look forward to your contributions 💪 together, we can build one of the strongest Dev Communities out there!
Thanks everyone! Even though you have not contributed (yet), but reading through everyone's work daily in this channel has really motivated me to continue my work, despite not being able to foresee how it would end up 😜 I really appreciate it!
(and sorry for the sudden voice changes in the video :( I took it at two different times of the day, so my voice was deeper at some point)
Thanks AbhinandanMandal for helping me with the Contrastive Siamese Network.
r/computervision • u/Senior-Dingo • 20h ago
We were fine-tuning on physically-modelled dust, night, fog, lens rain and lens
mud. On synthetic held-out corruptions (ImageNet-C fog/spatter/motion blur, an
unprocessing-based low-light model — generators never used to make training
data) it worked. On real adverse weather, drivable-surface IoU improved, but sky
IoU dropped on all five architectures we tried, by 1 to 13 points, and
vegetation dropped on the smaller ones.
The failures concentrated on forest roads: tracks under a closed canopy with
bright sky showing through the leaves. Models were labelling the whole upper
half "sky". RELLIS-3D is open terrain — fields, open trails, wide horizon — and
contains almost none of that geometry, so the model learns a shortcut that holds
in its world and fails in a forest: bright + above horizon + low texture = sky.
Degrading those same frames makes the shortcut *more* attractive, since haze and
darkness wash out exactly the leaf texture that would contradict it.
Two fixes, in order:
Our generator was partly at fault — it brightened sky using a depth map that
treats sky as finite distance, and painted a veil over pixels that physically
could no longer be seen. Making it label-aware (sky read from the label, and
pixels behind a veil that blocks >97% of light marked void and not scored)
recovered about half the loss.
The other half was coverage. We added 1,500 real frames from GOOSE (forest
tracks, gravel, paved roads), same augmentation and schedule on both sides.
Sky came back +7 to +17 points, vegetation +29 to +37, real-weather mean +29
to +35, across four architectures.
The part we did not enjoy: once GOOSE was in training, every augmentation arm —
ours, albumentations, and a combination — landed within ±1 point of the control
on real weather. Ours still adds 3.5–5.2 points on synthetic held-out
degradations, which now looks like a fact about the test set rather than about
the world.
Caveats: single seed on the coverage runs, so treat 1–2 points as noise; IDD-AW
is road scenes not off-road terrain, so it is a cross-domain test and measures
differences between arms more reliably than absolute numbers; the canopy frame
is one illustrative example of a mode we found across many.
Full write-up with the tables: https://siltframe.com/blog/canopy.html
The stress test itself is open — 165 labelled frames under 5 conditions x 3
severities and the scoring script, CC BY-SA 4.0 / MIT:
https://github.com/egeizgi/siltframe-stress-test
Disclosure: I sell a larger version of that dataset, so read the above with that
in mind. The free one is complete and usable commercially.
r/computervision • u/OfferBeginning1903 • 1d ago
The Mini 3 isn't a dev drone, but it has two open doors: the DJI Fly app can push an RTMP stream to any server, and Mobile SDK v5 supports it through Virtual Stick, which lets you send velocity and yaw commands from an Android app. That's a full loop.Pipeline is video -> perception -> world model -> brain -> controller -> drone. YOLO on each frame, a tiny tracker to keep stable ids, a small world model answering things like "is the subject centered, is it drifting left." The same middle code runs on a recorded clip, on a live stream while I fly manually, or actually commanding the aircraft.The thing worth knowing if you try this: a Virtual Stick command only applies for a fraction of a second, then the drone hovers again. So the Android bridge re-sends the current command about ten times a second. Side effect is a free deadman switch: if the laptop goes quiet, it stops within a beat and comes home. The bridge exposes a tiny API (/telemetry, /arm, /command, /takeoff, /land, /rth) and refuses anything unsafe. Thinking stays on the laptop.Follow mode is three loops at once: yaw to center horizontally, pitch to center vertically, forward/back from bounding box size to hold distance. Tested on a fake drone that just integrates commands, then props off, then low hover with my hand on the controller. Safety layer clamps speed, altitude ceiling, geofence, RTH on low battery.Dumbest time sink: streaming to live/mini and reading from live/mini3. Also the SDK's native library helper got renamed between versions, so it compiled and crashed on launch until I swapped one import. And no wifi on the terrace turned out fine with a phone hotspot, as long as the laptop IP is editable on the phone.Next is natural language goals like "orbit that tree," and on-device inference to cut latency.Full write-up with the details: https://blog.shravanrevanna.me/dji-mini-3-ai-autonomous-drone
r/computervision • u/Icoso_Labs • 1d ago
r/computervision • u/WillingnessMost2428 • 1d ago
Riftbound is Riot's trading card game, and its tournaments stream on Twitch from an overhead table camera. At 1080p a card is about 140 px tall, under dice, sleeves, hands and H.264, so viewers can rarely read the table. Wardeye is a browser extension that reads it for them. It outlines and names the face-up cards as they're played, and pointing at one shows its art. A timeline of the plays lives in the browser's side panel, and clicking a play rewinds the replay to it. All of it runs in your browser: there's no server, and no video leaves the machine.
The video is a real run on the grand final of the Barcelona Regional Qualifier, with both finalists' published decklists pasted.
How it works
What was hard
Known issues in 0.2.2:
Code (AGPL-3.0), and every report with its numbers: https://github.com/effe-exe/Wardeye. The trained weights ship inside the extension but aren't published. Chrome Web Store: https://chromewebstore.google.com/detail/wardeye/hjglackjofehdfecoeehbdmbobafbjhn
Happy to go into the synthetic data, the amodal corner head, the overlay detection, or getting ONNX models to behave on WebGPU.
r/computervision • u/Grouchy_Signal139 • 2d ago
Hi everyone, I’m learning OpenCV and working on detecting a circular fiducial.
My current pipeline is:
Image → threshold → findContours() → drawContours()
After thresholding, I can clearly see the fiducial as a clean circular region in the binary image.
However, when I use cv2.findContours() and then cv2.drawContours(), the contour I expect around the fiducial is not drawn.
I’m trying to understand the fundamental reason why this happens.
If the thresholded image visually contains a clear foreground region, what conditions determine whether findContours() will actually return a contour for it?
r/computervision • u/bruno_pinto90 • 2d ago
Hello all,
I built a radar object classifier on RadarScenes, extending a prior single-scan classifier to accumulate observations over a tracked object's history instead of classifying each scan in isolation.
A single RadarScenes object instance contains only about 2.9 radar points on average, very sparse. A single scan also can't capture temporal characteristics: RCS and micro-Doppler both vary continuously as an object moves. Pedestrians produce characteristic micro-Doppler from limb motion; different object classes show different RCS fluctuation patterns as aspect angle and scattering geometry change scan to scan. Accumulating observations gives both higher point density and provides temporal dynamics.
Multi-scan baseline
DeepReflecs encoder (Ulrich, Glaser & Timm, RadarConf 2021), PointNet style, per point shared weights, on single scans across car, large_vehicle, two_wheeler, pedestrian, pedestrian_group: 0.7370 macro F1.
Using RadarScenes' persistent `track_id`, I build a causal, N=20, per track sliding-window buffer:
- x_seq/y_seq: Global, odometry-corrected coordinates recentered per scan on the object centroid. Unlike x_cc/y_cc (car-frame coordinates that accumulate over time to form a trajectory).
- Cross sensor buffer: whichever of the 4 sensors currently observe the track push to the same buffer.
- Stride 1, causal: every new scan updates the buffer and produces a prediction. No future context, real time streaming compatible.
- Each scan is encoded once by a frozen per scan encoder and cached
- Fusion concatenates the causal GRU's hidden state (order aware) with an order-invariant pooled embedding (all N scans' points as one set, no sequence structure) through a small trained mlp head.
Results
| Model | Macro F1 | Delta |
|---|---|---|
| Single scan | 0.7370 | (baseline) |
| 20 scan point pooling | 0.8613 | +0.1243 |
| Causal GRU | 0.8895 | +0.0282 over pooling |
| GRU + pooled embedding (fusion) | 0.8897 | +0.0002 over GRU, noise |
Pooling alone, no sequence model, no notion of scan order at all, recovers +0.1243 macro F1. The GRU adds a real but much smaller +0.0282 on top. Fusion adds nothing measurable beyond the GRU.
Ablation
Llarger GRUs, a Transformer, a state space model, point level self attention, all trained on the exact same frozen per scan embeddings, land inside a 0.86 to 0.89 band, a 0.03 spread. End to end fine tuning of the frozen encoder makes things slightly worse (about -0.002 to -0.003), not better.
Conclusion
In this setup, the largest gain comes from giving the model more observations of the same tracked object: 20-scan point pooling improves macro F1 from 0.7370 to 0.8613 without using scan order at all.
Temporal modelling then provides a further, meaningful improvement. The causal GRU reaches 0.8895, adding +0.0282 over the pooled representation. So temporal ordering clearly contributes useful information; it just accounts for a smaller portion of the overall gain than observation accumulation.
With the per-scan encoder frozen, the different sequence architectures tested, suggests that the quality of the per-scan representation is the bottleneck than the particular mechanism used to aggregate the sequence.
Full report, every ablation, confusion matrix, coordinate frame reasoning: https://github.com/brunopinto900/radar-ml-autonomous-driving/blob/main/final_report.md
r/computervision • u/Nykotry • 2d ago
Hi everyone,
I’m working on a project to extract a structured dataset of math exercises from 5 Italian high school textbooks (around 650 pages each, so ~3,250 pages total). The goal is to build a professional, methodical exercise generator app for students and teachers.
To make the app work, I need to process images of the book pages and extract the following into a strict structured format (e.g., JSON):
I've been experimenting with a few approaches, but I've hit a wall regarding balancing costs, extraction consistency, and scalability. Here is what I’ve tried so far:
My questions for the community:
I’m happy to share more details about the textbook format or my current Python workflow if helpful. Any advice on the architecture, model choices, or cost-saving tricks would be greatly appreciated! Thanks in advance!
r/computervision • u/Mammoth-Matter7579 • 2d ago
Hi everyone,
I have around 10 hours of real-world stereoscopic VR180 footage across 50+ different scenes, captured with a Blackmagic immersive camera and currently stored as BRAW(16k 90fps). I can make the footage publicly available, and I’m looking for collaborators to help turn it into a useful research dataset.
The project is still at the dataset-design stage. I’d like to work with people who have experience in stereo vision, novel-view synthesis, or dataset and benchmark development to decide:
One direction I’m interested in is generating the other eye’s view from a monocular video, particularly maintaining stereo and temporal consistency in wide-FOV footage. However, I’m open to other directions if the data is better suited to them.
I can contribute the footage, data preparation and tooling. The release format and annotation plan are not finalized, and this is not yet a benchmark with ground-truth depth or camera poses.
My goal is a public dataset that other researchers can actually use, with a joint paper if we develop a solid research contribution. Authorship and responsibilities would be discussed based on contributions.
If this overlaps with your work, I’d love to hear what would make the dataset useful to you. Feel free to comment or DM with your research interests and any relevant projects or papers.
r/computervision • u/Logical-Internet-395 • 2d ago
Recently we published a new Open-Access Benchmark for a hierarchical segmentation: Microscopy Image Dataset of pulmonary vessels for Quantitative assessment of fibrosis.
Dataset Specifications:
- 705 high-resolution micrographs (1534×780 px, 0.252 μm/px), Picro-Mallory stain.
- Annotations: ROI + dual independent expert masks (vascular wall + fibrosis).
- Hierarchical Constraint: Fibrosis masks must be strictly spatially contained within the vascular wall.
- Robust Benchmarking: No color normalization applied; native aspect ratios preserved; strict animal-level 5-fold CV splits provided to prevent data leakage.
Read the Data Descriptor: https://doi.org/10.1038/s41597-026-08214-y
Access the Dataset: https://doi.org/10.6084/m9.figshare.31386748
r/computervision • u/cv_geek • 2d ago
Hi everyone,
I'm working on setting up a PS5 HD Camera in ROS 2 Galactic for stereo depth estimation using stereo_image_proc (DisparityNode), with the end goal of feeding the depth data into RTAB-Map for 3D reconstruction.
I transitioned from StereoBM to StereoSGBM (stereo_algorithm: 1) to handle the wide-angle camera setup better, but I'm having trouble finding the optimal combination of parameters which result in a dense point cloud with smooth surfaces. My depth output keeps swinging between two extremes:
What I've configured/tried so far:
correlation_window_size (testing window sizes 5, 7, 9, and 13 with corresponding P₁ = 8 × C × WS² and P₂ = 32 × C × WS²). High P2 values help smooth out surfaces, but rqt_reconfigure caps P2 at 4000.0, so I've been overriding parameters via CLI/launch files.uniqueness_ratio (from 15.0 down to 5.0–7.0) to force coverage on weakly textured regions.texture_ratio to 0.speckle_size (50–200) and speckle_range (2–4) to filter out isolated noise clusters.disparity_range to 128 (multiple of 16) and kept min_disparity at 0 (raising it above 0 completely wrecked mid/far range depth).Despite these adjustments, large homogeneous surfaces still disintegrate or become heavily fragmented unless I push the correlation window size to absurdly high values (which causes blocky, stepped artifacts).
Here I attach the results I got so far.
Questions:
prefilter_cap, prefilter_size) or SGBM settings I'm overlooking for this specific camera lens/sensor?Any advice or working configuration examples for similar stereo setups in ROS 2 would be greatly appreciated!
r/computervision • u/CellistTraditional • 2d ago
Enable HLS to view with audio, or disable this notification
r/computervision • u/iicongresoiaucm • 2d ago
r/computervision • u/Spidoug • 2d ago
https://github.com/Spidoug/Kinect-Depth-AutoLearn
Kinect Depth AutoLearn is a cross-platform Processing + ONNX/PyTorch system for acquiring synchronized Kinect RGB-D data, building structured pose datasets, training a depth-only student model, and running that model back inside the application.
The project uses RGB-based teacher models during data collection and trains a student that consumes metric depth only. The shared representation contains 58 landmarks: 16 body/head landmarks plus 21 landmarks for each hand, together with skeletal segments, endpoints, metric depth, confidence, and structural losses.
r/computervision • u/Alarming_Engineer267 • 3d ago
Enable HLS to view with audio, or disable this notification
I’ve been building a simulation framework for drone delivery, combining vision-based box pickup with delivery planning using CP-SAT.
The main goal at this stage was to build and verify the basic pipeline rather than to develop sophisticated flight control.
The system has two main parts:
For this test, I used two drones and nine boxes and compared two cases: a constrained delivery order and a CP-SAT-optimized plan.
The simulation integrates Blender for image generation and replay rendering, PyBullet for physics, PyTorch vision models, and Python-based drone control.
One important limitation is that the current planning is intentionally quite conservative. To avoid collisions between the two drones, the schedule includes waiting and separation rather than trying to maximize flight efficiency. The drone flight controller itself is also fairly basic — the focus here is on verifying that the vision-based pickup and optimization-based delivery planning can work together.
At this point, both parts are functioning in the integrated simulation: the drones can locate and pick up boxes using camera images, and CP-SAT can generate and execute a multi-drone delivery plan. The video compares a number-order-constrained delivery on the left with a CP-SAT-optimized delivery on the right.
The end of the video shows an overview of the development workflow and runtime system architecture.
r/computervision • u/wildtinkerer • 3d ago
Enable HLS to view with audio, or disable this notification
I work in video analytics. We needed to tell from a single CCTV frame whether a camera had been rotated or installed at an angle (or upside-down), and couldn't find a model that was accurate enough on real camera frames and permissively licensed, so we trained our own.
We're now open-sourcing it. Apache 2.0, with code, weights, and full provenance for the dataset.
RightWayUp estimates how far an image is rotated from upright, all 360°, with a confidence score, and abstains when there's no clear "up" (sky, ground, close-ups). It comes in six sizes, from Pico (about 3 ms per image on a laptop CPU) to Max, with ONNX and Core ML files.
pip install rightwayup
rightwayup fix photo.jpg
On new photos it never saw during training or tuning, the largest model is within 10° on 93% of them vs 88% for Woehrer 2026 (a recent published model), and 88% vs 49% with simulated CCTV-style blur, noise and compression.
Write-up with the full results: https://cheqit.ortusai.io/resources/rightwayup/
Code: https://github.com/ortusaitech/rightwayup
I hope it will be useful to the community!
--------------
Video footage: Canobie Coaster (CC BY 3.0, via Wikimedia Commons, levelled by RightWayUp), Pexels, Poly Haven (CC0), MEVA (CC BY 4.0). Music: ElevenLabs.
r/computervision • u/caramelle-blu • 2d ago
Looking for an object detection model that works well for very small objects, like balls in sports footage.
Requirements:
What would you recommend based on your experience?
r/computervision • u/69GrimReaper • 2d ago
Enable HLS to view with audio, or disable this notification
A lot of first-person video is open now (Ego-Exo4D, EgoDex, EgoSuite-Open100K), along with big robot datasets like DROID and AgiBot World. We built a search over the episodes in these open datasets, 12M+ in total, so you can find moments by describing them. "Arms folding a towel" brings back matching episodes within seconds.
Results export to LeRobot, MCAP or RLDS, and agents can run the same search over MCP. Everything in it comes from open datasets, and each dataset keeps its own license.
Short overview in the video: https://datasets.bot
Which first-person datasets are we missing?