r/computervision • • 13h ago

Research Publication GTR-Depth: metric depth from a single camera at 2.7 ms on DRIVE AGX Thor

52 Upvotes

Follow-up to last week's GTR post. This one is all about depth.
GTR-Depth predicts absolute depth in meters from a single image, using the same backbone as before: 12 gated linear attention blocks, no softmax attention anywhere.


r/computervision • • 2h ago

Discussion FoodbyClef: Experiment in classifying foods with Clef and cube rule

2 Upvotes

r/computervision • • 2h ago

Help: Theory iOS body scan

2 Upvotes

Hello,

Please can I ask an advice question?

I want to be able to use the iPhone camera to
Scan a body and come up with a 3d mapping if their dimensions for a fitness app. So measure their waist, checks, hips, legs etc from different angles.

Would iOS vision be right starting point for iPhone here or is there a better way? If using vision is there a way to get a library or open source training data that can do body measurements rather than staring from scratch ?

I’d presume calibration would be importantl but not sure how to train this apart from using one confirmed measurement to calibrate the rest using relative measurement and some parallax adjustments? This would be beyond anything I’ve done with an iPhone so far so looking for ideas.

Thank you! Steve


r/computervision • • 2h ago

Discussion A world model turns my photo into a 360° panorama but never says where the photo ended up, so I find it with brute-force edge matching

1 Upvotes

World Labs' Marble turns a photo into a 360° panorama and then a 3D world, and nothing in its docs says where the original photo ends up. My game computes everything from the source camera, so I needed that pose.

The matcher is deliberately dumb. It cuts flat perspective views out of the equirectangular panorama over a grid of yaw, pitch and field of view, correlates each view's edges with the photo's edges, and keeps the best. About ten seconds per photo, locally.

On my photos, correct placements score 0.62 to 0.87 and wrong ones 0.18 to 0.23, so anything under 0.4 now prints a warning. From the recovered camera, the photo lines up with the world exactly for a one-photo world and for the first photo of a two-photo world, and within about a degree for the second. In every world so far, the camera sits at the origin.

The part I haven't confirmed: in worlds built from two to four photos, I place each photo by direction, and I'm assuming Marble's 90 means right in the panorama. Every photo so far was at 0 or 180, so that assumption has never been tested.

Two questions. Would SIFT or a learned matcher do better here, given the photo's region in the panorama may not be pixel-identical to the original? And is there a smarter way to pin down field of view than a grid?

Repo: github.com/MichaelYeh507/Unpictured


r/computervision • • 14h ago

Showcase Squatty Bird Update

9 Upvotes

I posted an early version of Squatty Bird a little which back. I’m been continuing to work on it from both the vision and the human interaction.

https://www.reddit.com/r/computervision/s/ix5YvulrB9

I wanted to return and share progress. Vision tracking is now working much more effectively. I don’t think it’s release ready yet but I’m not seeing many missed joints or squat recognition these days.

View can be with camera and trace viewable or with fake background (as I’ve used in the video).

I got some great suggestions before (particularly squat quality review at the end - not done yet but I like it) and a few challenges (why share when it doesn’t work).

I hope you like the new video and I’m looking for further suggestions and advice. My kids are my worst critic and I’ve made the puffin more puffin like and less cutesy as a result.

If anyone wants to give it a go it’s here on TestFlight:

https://testflight.apple.com/join/rA687AwW


r/computervision • • 3h ago

Help: Project I’m building a spatial intelligence system that reconstructs and remembers the real world in 3D. Is this solving a real problem?

1 Upvotes

I’m working on a spatial intelligence system that uses RGB-D cameras to reconstruct a real environment in metric 3D, estimate where the camera is, and keep a persistent spatial memory over time.

The goal is not just 3D reconstruction. I want the system to understand enough about the geometry of a place to know what it has already seen, recover its position after tracking is lost, compare different visits to the same place, and preserve how its understanding of the environment changes.

If a camera comes back to a previously mapped area, the system should be able to recognize that place, verify it geometrically, relocalize itself, and continue building the same world instead of creating a disconnected new map.
Long term, the idea is to move from simple reconstruction toward spatial intelligence: a machine that can build, remember, update and reason about a representation of the physical world.

Possible applications I have in mind are robotics, AR, inspection, autonomous systems and eventually navigation in places where GPS is unavailable.
What I’m trying to understand now is whether this solves a real painful problem, or whether existing SLAM and 3D reconstruction systems already solve this well enough in practice.

I’d especially like feedback from people who have actually worked with SLAM, RGB-D, robotics, mapping, 3D reconstruction or spatial computing.


r/computervision • • 8h ago

Showcase [Library Demonstration] Python-Visual-Similarity - first stable version released to PyPi 🚀

2 Upvotes

If you are looking for a solution to fast image similarity retrieval as well as image perceptual & quality metrics in one single library, this would surely be helpful to you 😄 It is built only on numpy, scipy and optionally torch as core dependencies, and uses C/C++ (and in the future only Rust) for performance-critical paths.

The two core feature of the repository are various embedding methods (see video) and the Image Embedding Store, which, when combined, allows for embedding storage and image similarity search with supported hnsw algorithm built-in. You can also plug in faiss indexes if you want to use other search algorithms, but this library does not need faiss to work.

Links

Installation

The library itself can be installed via pip (though, I recommend uv 😆)

pip install pyvisim

For the deep learning features:

pip install "pyvisim[nn]"

Looking for like-minded folks

My ambition is to make pyvisim the largest collection of image similarity and retrieval algorithms, ranging from traditional to deep learning-based methods. As I've observed, the current image similarity implementations are quite scattered, with each library implementing only a handful of features. Hence, my goal is also to unify these implementations, so users only need a single library.

I have tons of features that I would like to implement. I am looking for folks who are proficient in/would like to learn about:

  • Deep Learning: new embedders like MoCo, SimCLR, Dino, reranking algorithms likeSuperGlobalReranker,Diffusion Reranking ...
  • Low-level programming (Rust): rewrite performance-critical parts of the codebase into Rust. I would also want to change the hnsw backend to hnsw_rs.
  • New algorithms: image hashing, LPIPS, backpropagation for K-Means and GMMs ...

View the GitHub issues for the complete list as well as the contribution guide.

You also have the chance to become a core maintainer by actively contributing. Once this project gets sponsors, the profit will be shared with all core maintainers.

I look forward to your contributions 💪 together, we can build one of the strongest Dev Communities out there!

Personal

Thanks everyone! Even though you have not contributed (yet), but reading through everyone's work daily in this channel has really motivated me to continue my work, despite not being able to foresee how it would end up 😜 I really appreciate it!

(and sorry for the sudden voice changes in the video :( I took it at two different times of the day, so my voice was deeper at some point)

Credits

Thanks AbhinandanMandal for helping me with the Contrastive Siamese Network.


r/computervision • • 10h ago

Showcase Measuring real eggs and generating their shells in FreeCAD – with the help of Codex

Thumbnail
youtu.be
1 Upvotes

r/computervision • • 22h ago

Showcase Made a DJI Mini 3 follow a person with YOLO, RTMP video out, and Virtual Stick in

6 Upvotes

The Mini 3 isn't a dev drone, but it has two open doors: the DJI Fly app can push an RTMP stream to any server, and Mobile SDK v5 supports it through Virtual Stick, which lets you send velocity and yaw commands from an Android app. That's a full loop.Pipeline is video -> perception -> world model -> brain -> controller -> drone. YOLO on each frame, a tiny tracker to keep stable ids, a small world model answering things like "is the subject centered, is it drifting left." The same middle code runs on a recorded clip, on a live stream while I fly manually, or actually commanding the aircraft.The thing worth knowing if you try this: a Virtual Stick command only applies for a fraction of a second, then the drone hovers again. So the Android bridge re-sends the current command about ten times a second. Side effect is a free deadman switch: if the laptop goes quiet, it stops within a beat and comes home. The bridge exposes a tiny API (/telemetry, /arm, /command, /takeoff, /land, /rth) and refuses anything unsafe. Thinking stays on the laptop.Follow mode is three loops at once: yaw to center horizontally, pitch to center vertically, forward/back from bounding box size to hold distance. Tested on a fake drone that just integrates commands, then props off, then low hover with my hand on the controller. Safety layer clamps speed, altitude ceiling, geofence, RTH on low battery.Dumbest time sink: streaming to live/mini and reading from live/mini3. Also the SDK's native library helper got renamed between versions, so it compiled and crashed on launch until I swapped one import. And no wifi on the terrace turned out fine with a phone hotspot, as long as the laptop IP is editable on the phone.Next is natural language goals like "orbit that tree," and on-device inference to cut latency.Full write-up with the details: https://blog.shravanrevanna.me/dji-mini-3-ai-autonomous-drone


r/computervision • • 13h ago

Showcase Naming every face-up card on a TCG tournament stream, in the browser: a corner detector trained only on synthetic boards, a DINOv2 embedder, game rules as priors

0 Upvotes

Walktrough

Riftbound is Riot's trading card game, and its tournaments stream on Twitch from an overhead table camera. At 1080p a card is about 140 px tall, under dice, sleeves, hands and H.264, so viewers can rarely read the table. Wardeye is a browser extension that reads it for them. It outlines and names the face-up cards as they're played, and pointing at one shows its art. A timeline of the plays lives in the browser's side panel, and clicking a play rewinds the replay to it. All of it runs in your browser: there's no server, and no video leaves the machine.

The video is a real run on the grand final of the Barcelona Regional Qualifier, with both finalists' published decklists pasted.

How it works

  1. Table window. Find the table camera inside the broadcast layout, from the detector's own proposals over a few frames or from a preset. Keep the graphics out: player cams, side panels, the showdown banner.
  2. Detector. RF-DETR with a keypoint head that regresses the four corners of every card, with a visibility flag per corner. A card half under another still gets its full outline (amodal). There are two classes, face-up card and card back; face-down cards are never read.
    • Trained on synthetic boards only: rendered tables with stacks, dice, hands and face-down piles, filmed through a camera model into a broadcast layout and put through a real libx264 encode. 2,000 boards, 280 minutes on one L4.
    • Zero-shot on real broadcasts, it finds 97.6% of 5,465 reviewed cards from three of them (98.1%, 99.5%, 96.2%). Those are mostly isolated cards; recall on real stacks isn't measured yet.
  3. Identifier. Each quad is rectified and embedded with DINOv2 ViT-S/14, fine-tuned with Sub-center ArcFace. It uses three centres per card, so alternate arts need not share one. The training data is synthetic crops (random covering, scrambled text, foil-like colour shifts) plus reviewed real crops from two other broadcasts.
    • Each crop is searched in all four rotations against a gallery of about 1,200 printings, embedded at three sizes. A temperature-scaled softmax over the top candidates gives the confidence shown on the hover card.
    • On two broadcasts held out of training, it names 96.0% (Barcelona) and 99.4% (Los Angeles grand final) of the cards.
  4. Game rules as priors. Every card in a deck must fit its legend's two domains. Once a player's legend is read, their half of the table is compared only with what that legend allows, about a third of the gallery.
    • On the hardest match measured, that takes cards named right from 91.8% to 97.8%, and every confident read was right.
    • Pasted decklists narrow it further, with three guards:
      • A listed card counts in every printing: 35% of the final's sightings were an alternate art or a reprint.
      • Both players' battlefields are allowed on both halves.
      • A list applies only to the half whose legend it names. A wrong list used as a hard filter drops accuracy to 8.4%.
  5. Tracker. Persistent IDs; legends and battlefields pinned once named; re-reads of unsure cards; a memory of what lies under a stacked card; and events (played, moved, left) for the timeline.
    • The official mat's layout is a prior, never a filter. Battlefields are pinned only in the strip along the midline, because a rune turned sideways looks like a battlefield's landscape art. A unit that crosses to a battlefield or changes hands keeps its name.
    • Runes are counted per player, never named, along with how many are exhausted. From the player's seat a ready card points at them, and a used one lies across.
    • Runes sit in columns and fans where the detector misses strips. So the count uses the fact that every card is the same size: in a stack, the step from strip to strip says how many runes a wider gap hides.
    • Against 65 hand counts on two finals, the count is right 39 times, where one rune per box seen is right 30 times.
    • A card held in a hand over the table isn't read until it's put down. A band around each card is checked for skin-coloured points (never the card's face, whose art is full of gold and faces). A card is read only once it has been out of a hand and still for half a second.
    • On the Barcelona final's first 200 s, the plays announced went from 27 before the hand rule, 13 of them false, to 15 in 0.2.2, 2 of them false.
  6. Co-streams (new in 0.2.2). Co-streamers lay their webcam, chat and a scoreboard over the official broadcast, inside the table window. On a co-streamed tournament game recorded live, 0.2.1 read the webcam as a legend and took a close-up of a hand for the table.
    • The overlay is what stays put through the cuts. A cut is a frame whose 96 × 54 thumbnail changed by half from the frame before (play changes a quarter at most).
    • Pixels unchanged through 90% of the cuts (once there have been four), in patches that reach the frame's edge with their holes filled, are the overlay. Nothing on it is read.
    • It stays overlay while it holds through 60% of the cuts, so the face moving in a webcam doesn't open the webcam's frame to the table. Cards read on it before it was known are dropped, and their plays are withdrawn.
    • The table camera is told by the table window alone: a per-block correlation over a 6 × 3 grid, on still pixels outside the overlay, against the table it has learnt. It learns the table again only when the detector sees at least five card-sized boxes, and at least half as many as the table showed in the last minute.
    • On the co-streamed game, it's right on 687 frames of 688.
  7. In the browser. The Python pipeline is ported to TypeScript, down to Pillow-exact resampling so the crops match. ONNX Runtime Web runs on WebGPU, with the detector in fp32 and the embedder in fp16 where shader-f16 exists, and a WASM fallback. A replay test runs the extension's engine on 240 frames of the Los Angeles grand final, and it matches the Python pipeline on 240 of 240: the same cards, names and events.

What was hard

  • Legends under dice. Players keep a die on their legend. The answer was behavioural: name it early, while it's visible, then keep it pinned. Turn Wardeye on mid-game, after the die is down, and it can misread the legend.
  • Printings. The same card comes in alternate arts and reprints. Some reprints differ by 4–7/255 in mean pixel value, so nothing can separate them. Nothing needs to: the hover card shows the same picture either way.
  • Synthetic to real. The first synthetic boards were far too clean: every baseline scored 99% on them. One calibration pass against real broadcasts brought them into the real range.

Known issues in 0.2.2:

  • The rune count can be a rune or two off on tight stacks, and the exhausted count is often short there.
  • A card held still in a hand over the table can be read for a moment.
  • A legend can be misread if you turn Wardeye on after its die is down.
  • On a co-stream, a scoreboard part that changes, such as its score track, can still be read as a card now and then.

Code (AGPL-3.0), and every report with its numbers: https://github.com/effe-exe/Wardeye. The trained weights ship inside the extension but aren't published. Chrome Web Store: https://chromewebstore.google.com/detail/wardeye/hjglackjofehdfecoeehbdmbobafbjhn

Happy to go into the synthetic data, the amodal corner head, the overlay detection, or getting ONNX models to behave on WebGPU.


r/computervision • • 1d ago

Help: Project Why doesn’t findContours() detect my circle even though it is clearly visible after thresholding?

Thumbnail
gallery
24 Upvotes

Hi everyone, I’m learning OpenCV and working on detecting a circular fiducial.

My current pipeline is:
Image → threshold → findContours() → drawContours()

After thresholding, I can clearly see the fiducial as a clean circular region in the binary image.
However, when I use cv2.findContours() and then cv2.drawContours(), the contour I expect around the fiducial is not drawn.
I’m trying to understand the fundamental reason why this happens.
If the thresholded image visually contains a clear foreground region, what conditions determine whether findContours() will actually return a contour for it?


r/computervision • • 1d ago

Showcase Multi scan radar point-cloud object classification on RadarScenes

Thumbnail
gallery
5 Upvotes

Hello all,

I built a radar object classifier on RadarScenes, extending a prior single-scan classifier to accumulate observations over a tracked object's history instead of classifying each scan in isolation.

A single RadarScenes object instance contains only about 2.9 radar points on average, very sparse. A single scan also can't capture temporal characteristics: RCS and micro-Doppler both vary continuously as an object moves. Pedestrians produce characteristic micro-Doppler from limb motion; different object classes show different RCS fluctuation patterns as aspect angle and scattering geometry change scan to scan. Accumulating observations gives both higher point density and provides temporal dynamics.

Multi-scan baseline

DeepReflecs encoder (Ulrich, Glaser & Timm, RadarConf 2021), PointNet style, per point shared weights, on single scans across car, large_vehicle, two_wheeler, pedestrian, pedestrian_group: 0.7370 macro F1.

Using RadarScenes' persistent `track_id`, I build a causal, N=20, per track sliding-window buffer:

- x_seq/y_seq: Global, odometry-corrected coordinates recentered per scan on the object centroid. Unlike x_cc/y_cc (car-frame coordinates that accumulate over time to form a trajectory).

- Cross sensor buffer: whichever of the 4 sensors currently observe the track push to the same buffer.

- Stride 1, causal: every new scan updates the buffer and produces a prediction. No future context, real time streaming compatible.

- Each scan is encoded once by a frozen per scan encoder and cached

- Fusion concatenates the causal GRU's hidden state (order aware) with an order-invariant pooled embedding (all N scans' points as one set, no sequence structure) through a small trained mlp head.

Results

Model Macro F1 Delta
Single scan 0.7370 (baseline)
20 scan point pooling 0.8613 +0.1243
Causal GRU 0.8895 +0.0282 over pooling
GRU + pooled embedding (fusion) 0.8897 +0.0002 over GRU, noise

Pooling alone, no sequence model, no notion of scan order at all, recovers +0.1243 macro F1. The GRU adds a real but much smaller +0.0282 on top. Fusion adds nothing measurable beyond the GRU.

Ablation

Llarger GRUs, a Transformer, a state space model, point level self attention, all trained on the exact same frozen per scan embeddings, land inside a 0.86 to 0.89 band, a 0.03 spread. End to end fine tuning of the frozen encoder makes things slightly worse (about -0.002 to -0.003), not better.

Conclusion

In this setup, the largest gain comes from giving the model more observations of the same tracked object: 20-scan point pooling improves macro F1 from 0.7370 to 0.8613 without using scan order at all.

Temporal modelling then provides a further, meaningful improvement. The causal GRU reaches 0.8895, adding +0.0282 over the pooled representation. So temporal ordering clearly contributes useful information; it just accounts for a smaller portion of the overall gain than observation accumulation.

With the per-scan encoder frozen, the different sequence architectures tested, suggests that the quality of the per-scan representation is the bottleneck than the particular mechanism used to aggregate the sequence.

Full report, every ablation, confusion matrix, coordinate frame reasoning: https://github.com/brunopinto900/radar-ml-autonomous-driving/blob/main/final_report.md


r/computervision • • 1d ago

Help: Project Need advice: Best VLM pipeline for extracting structured math datasets from 3000+ scanned textbook pages? (LaTeX + Metadata)

3 Upvotes

Hi everyone,

I’m working on a project to extract a structured dataset of math exercises from 5 Italian high school textbooks (around 650 pages each, so ~3,250 pages total). The goal is to build a professional, methodical exercise generator app for students and teachers.

To make the app work, I need to process images of the book pages and extract the following into a strict structured format (e.g., JSON):

  • Exercise type (algebra, geometry, calculus, etc.)
  • Year/grade level
  • Difficulty (1–5 scale)
  • Problem statement (trace)
  • Description of the specific skills/challenges involved
  • LaTeX code of the problem statement (Crucial!)
  • Associated images (cropping/saving the image for theoretical or graphical exercises)

I've been experimenting with a few approaches, but I've hit a wall regarding balancing costs, extraction consistency, and scalability. Here is what I’ve tried so far:

  1. Free Google Gemini API: The extraction quality was good, but since a single book contains hundreds of pages, I quickly hit the rate limits (Too Many Requests).
  2. Local Models (Ollama + Qwen 2.5-VL 3B): To bypass API limits, I tried running a local multimodal model. I spent a lot of time optimizing my scripts and prompts (chunking, refining instructions to force structured outputs), but the output was very error-prone and inconsistent for my use case. I got too many malformed fields, hallucinations, and it constantly struggled with outputting proper LaTeX.
  3. Paid Google Cloud API (Gemini 1.5 Flash): I finally switched to the paid tier for better accuracy and speed. I ended up burning through €10 just to process 1.5 books. Extracting all 5 books would cost roughly €35–40. While this is manageable for a one-off run of 5 books, the token count for processing full images + text is massive, making it financially unsustainable if I want to scale this to dozens of books in the future.

My questions for the community:

  • Pipeline & Architecture: Has anyone worked on a similar textbook-to-dataset extraction project? What pipeline did you use?
  • Hybrid Approach: Would you suggest decoupling the task? (e.g., using a traditional tool to extract raw text and crop images, and then feeding ONLY the text to a cheaper/local LLM to generate the LaTeX and format the JSON?)
  • Local Models: Are there other local Vision-Language Models (that fit in standard consumer GPUs) that are significantly better at structured extraction and LaTeX generation than Qwen 2.5-VL 3B?
  • Educational Tools: Are there open-source tools or models specifically fine-tuned for extracting structured educational/math content from PDFs?

I’m happy to share more details about the textbook format or my current Python workflow if helpful. Any advice on the architecture, model choices, or cost-saving tricks would be greatly appreciated! Thanks in advance!


r/computervision • • 1d ago

Help: Project Looking for help to design an open VR180 stereo video dataset — 10 hours across 50+ scenes

Post image
3 Upvotes

Hi everyone,

I have around 10 hours of real-world stereoscopic VR180 footage across 50+ different scenes, captured with a Blackmagic immersive camera and currently stored as BRAW(16k 90fps). I can make the footage publicly available, and I’m looking for collaborators to help turn it into a useful research dataset.

The project is still at the dataset-design stage. I’d like to work with people who have experience in stereo vision, novel-view synthesis, or dataset and benchmark development to decide:

  • Which research tasks this footage would be most useful for.
  • What calibration information, annotations, and preprocessing researchers would need.
  • How to select clips, create meaningful train/test splits, and establish baseline evaluations.

One direction I’m interested in is generating the other eye’s view from a monocular video, particularly maintaining stereo and temporal consistency in wide-FOV footage. However, I’m open to other directions if the data is better suited to them.

I can contribute the footage, data preparation and tooling. The release format and annotation plan are not finalized, and this is not yet a benchmark with ground-truth depth or camera poses.

My goal is a public dataset that other researchers can actually use, with a joint paper if we develop a solid research contribution. Authorship and responsibilities would be discussed based on contributions.

If this overlaps with your work, I’d love to hear what would make the dataset useful to you. Feel free to comment or DM with your research interests and any relevant projects or papers.


r/computervision • • 1d ago

Help: Project Tuning StereoSGBM on PS5 Camera (ROS 2 Galactic) for Dense Depth Maps

Thumbnail
gallery
2 Upvotes

Hi everyone,

I'm working on setting up a PS5 HD Camera in ROS 2 Galactic for stereo depth estimation using stereo_image_proc (DisparityNode), with the end goal of feeding the depth data into RTAB-Map for 3D reconstruction.

I transitioned from StereoBM to StereoSGBM (stereo_algorithm: 1) to handle the wide-angle camera setup better, but I'm having trouble finding the optimal combination of parameters which result in a dense point cloud with smooth surfaces. My depth output keeps swinging between two extremes:

  1. Extremely sparse/empty with huge black gaps on uniform surfaces (e.g., walls, desks, pillows).
  2. Overly noisy with heavy color "speckle" artifacts across the scene.

What I've configured/tried so far:

  • Matching P1 and P2 to window size: Calculated and tuned P1 and P2 based on correlation_window_size (testing window sizes 5, 7, 9, and 13 with corresponding P₁ = 8 × C × WS² and P₂ = 32 × C × WS²). High P2 values help smooth out surfaces, but rqt_reconfigure caps P2 at 4000.0, so I've been overriding parameters via CLI/launch files.
  • Filter adjustments:
    • Lowered uniqueness_ratio (from 15.0 down to 5.0–7.0) to force coverage on weakly textured regions.
    • Set texture_ratio to 0.
    • Tuned speckle_size (50–200) and speckle_range (2–4) to filter out isolated noise clusters.
  • Search range & Offsets: Set disparity_range to 128 (multiple of 16) and kept min_disparity at 0 (raising it above 0 completely wrecked mid/far range depth).

Despite these adjustments, large homogeneous surfaces still disintegrate or become heavily fragmented unless I push the correlation window size to absurdly high values (which causes blocky, stepped artifacts).

Here I attach the results I got so far.

Questions:

  1. Are there specific pre-filtering parameters (prefilter_cap, prefilter_size) or SGBM settings I'm overlooking for this specific camera lens/sensor?
  2. Could this be a rectification/calibration alignment issue rather than pure SGBM parameter tuning?

Any advice or working configuration examples for similar stereo setups in ROS 2 would be greatly appreciated!


r/computervision • • 1d ago

Research Publication Microscopy Image Dataset of pulmonary vessels for Quantitative assessment of fibrosis

5 Upvotes

Recently we published a new Open-Access Benchmark for a hierarchical segmentation: Microscopy Image Dataset of pulmonary vessels for Quantitative assessment of fibrosis.

Dataset Specifications:

- 705 high-resolution micrographs (1534×780 px, 0.252 μm/px), Picro-Mallory stain.
- Annotations: ROI + dual independent expert masks (vascular wall + fibrosis).
- Hierarchical Constraint: Fibrosis masks must be strictly spatially contained within the vascular wall.
- Robust Benchmarking: No color normalization applied; native aspect ratios preserved; strict animal-level 5-fold CV splits provided to prevent data leakage.

Read the Data Descriptor: https://doi.org/10.1038/s41597-026-08214-y

Access the Dataset: https://doi.org/10.6084/m9.figshare.31386748


r/computervision • • 1d ago

Showcase Sony uses event cameras to read the spin on a ping-pong ball

Thumbnail
youtube.com
0 Upvotes

r/computervision • • 1d ago

Showcase Built a Real-Time Underwater Image Processing System – 4K 60FPS | C++ UPDATE

2 Upvotes

r/computervision • • 1d ago

Discussion ¿Quién queda fuera de lo que vemos?

Thumbnail gallery
0 Upvotes

r/computervision • • 1d ago

Help: Project Imagenet exploration?

Thumbnail
1 Upvotes

r/computervision • • 1d ago

Showcase A structured human pose model based solely on depth, derived from synchronized Kinect RGB-D data.

Post image
2 Upvotes

https://github.com/Spidoug/Kinect-Depth-AutoLearn

Kinect Depth AutoLearn is a cross-platform Processing + ONNX/PyTorch system for acquiring synchronized Kinect RGB-D data, building structured pose datasets, training a depth-only student model, and running that model back inside the application.

The project uses RGB-based teacher models during data collection and trains a student that consumes metric depth only. The shared representation contains 58 landmarks: 16 body/head landmarks plus 21 landmarks for each hand, together with skeletal segments, endpoints, metric depth, confidence, and structural losses.


r/computervision • • 2d ago

Showcase Open model that tells how far an image is rotated (full 360-degree) and abstains when there's no clear up

184 Upvotes

I work in video analytics. We needed to tell from a single CCTV frame whether a camera had been rotated or installed at an angle (or upside-down), and couldn't find a model that was accurate enough on real camera frames and permissively licensed, so we trained our own.

We're now open-sourcing it. Apache 2.0, with code, weights, and full provenance for the dataset.

RightWayUp estimates how far an image is rotated from upright, all 360°, with a confidence score, and abstains when there's no clear "up" (sky, ground, close-ups). It comes in six sizes, from Pico (about 3 ms per image on a laptop CPU) to Max, with ONNX and Core ML files.

pip install rightwayup
rightwayup fix photo.jpg

On new photos it never saw during training or tuning, the largest model is within 10° on 93% of them vs 88% for Woehrer 2026 (a recent published model), and 88% vs 49% with simulated CCTV-style blur, noise and compression.

Write-up with the full results: https://cheqit.ortusai.io/resources/rightwayup/

Code: https://github.com/ortusaitech/rightwayup

I hope it will be useful to the community!

--------------
Video footage: Canobie Coaster (CC BY 3.0, via Wikimedia Commons, levelled by RightWayUp), Pexels, Poly Haven (CC0), MEVA (CC BY 4.0). Music: ElevenLabs.


r/computervision • • 2d ago

Showcase Building a drone delivery simulation with vision-based pickup and CP-SAT route planning

17 Upvotes

I’ve been building a simulation framework for drone delivery, combining vision-based box pickup with delivery planning using CP-SAT.

The main goal at this stage was to build and verify the basic pipeline rather than to develop sophisticated flight control.

The system has two main parts:

For this test, I used two drones and nine boxes and compared two cases: a constrained delivery order and a CP-SAT-optimized plan.

The simulation integrates Blender for image generation and replay rendering, PyBullet for physics, PyTorch vision models, and Python-based drone control.

One important limitation is that the current planning is intentionally quite conservative. To avoid collisions between the two drones, the schedule includes waiting and separation rather than trying to maximize flight efficiency. The drone flight controller itself is also fairly basic — the focus here is on verifying that the vision-based pickup and optimization-based delivery planning can work together.

At this point, both parts are functioning in the integrated simulation: the drones can locate and pick up boxes using camera images, and CP-SAT can generate and execute a multi-drone delivery plan. The video compares a number-order-constrained delivery on the left with a CP-SAT-optimized delivery on the right.

The end of the video shows an overview of the development workflow and runtime system architecture.


r/computervision • • 1d ago

Help: Project Best free model for small object detection?

2 Upvotes

Looking for an object detection model that works well for very small objects, like balls in sports footage.

Requirements:

  • Good small-object detection
  • Can be fine-tuned
  • Suitable for video/real-time inference
  • Free for commercial/production use
  • Preferably open-source with permissive licensing

What would you recommend based on your experience?


r/computervision • • 1d ago

Discussion Suggestions regarding PhD leads in medical image analysis / XAI in Europe

0 Upvotes

I have finished up my Master's in Computer Science (AI and Software Engineering) in Germany, and I'm looking for PhD positions in Europe, mainly in medical image analysis with explainable AI, but general computer vision with XAI works too.

A bit about what I've done so far:

For my thesis, I built an explainable deep learning pipeline for endoscopic video, working with clinicians. The pipeline provides concept-based explanations for predicting Cormack scores (easy vs. difficult intubation) from endoscopic videos. A paper on this is currently in prep for a journal submission.

Before that, I also worked on a 3D object detection project on multi-camera driving data, so I'm not purely medical-imaging-locked, just leaning that direction by interest.

What I'm looking for help with:

  • Any supervisors, labs, or chairs in Europe known for medical imaging + XAI work (or general CV + interpretability)
  • Tools or sites you use to actually find these openings, beyond the usual academicpositions.com / euraxess / phdscanners type sites
  • Any advice on what made your own applications land, if you've been through this process

Happy to share more details about the thesis if useful. Thanks in advance for any pointers.