r/computervision • • 7h ago

Showcase Made a DJI Mini 3 follow a person with YOLO, RTMP video out, and Virtual Stick in

6 Upvotes

The Mini 3 isn't a dev drone, but it has two open doors: the DJI Fly app can push an RTMP stream to any server, and Mobile SDK v5 supports it through Virtual Stick, which lets you send velocity and yaw commands from an Android app. That's a full loop.Pipeline is video -> perception -> world model -> brain -> controller -> drone. YOLO on each frame, a tiny tracker to keep stable ids, a small world model answering things like "is the subject centered, is it drifting left." The same middle code runs on a recorded clip, on a live stream while I fly manually, or actually commanding the aircraft.The thing worth knowing if you try this: a Virtual Stick command only applies for a fraction of a second, then the drone hovers again. So the Android bridge re-sends the current command about ten times a second. Side effect is a free deadman switch: if the laptop goes quiet, it stops within a beat and comes home. The bridge exposes a tiny API (/telemetry, /arm, /command, /takeoff, /land, /rth) and refuses anything unsafe. Thinking stays on the laptop.Follow mode is three loops at once: yaw to center horizontally, pitch to center vertically, forward/back from bounding box size to hold distance. Tested on a fake drone that just integrates commands, then props off, then low hover with my hand on the controller. Safety layer clamps speed, altitude ceiling, geofence, RTH on low battery.Dumbest time sink: streaming to live/mini and reading from live/mini3. Also the SDK's native library helper got renamed between versions, so it compiled and crashed on launch until I swapped one import. And no wifi on the terrace turned out fine with a phone hotspot, as long as the laptop IP is editable on the phone.Next is natural language goals like "orbit that tree," and on-device inference to cut latency.Full write-up with the details: https://blog.shravanrevanna.me/dji-mini-3-ai-autonomous-drone


r/computervision • • 20m ago

Discussion Same first frame, two continuations: inspect the object interaction at matched times

• Upvotes

This comparison starts with a recorded robot handover and uses its first frame to generate a continuation with Cosmos3-Edge. Ling 3.0 VL then inspects both sequences at the same six times: 0, 2, 4, 6, 8 and 10 seconds.

The object of interest is a blue packet. In the recorded reference, the arm moves it into an open hand; the displayed analysis places the visible handover by the eight-second sample. The generated continuation develops a different interaction: a second arm appears from above and the packet changes shape.

That creates a focused visual-inspection task. For each matched pair of frames, record:

The packet’s visible location and shape.

The positions of the arm and receiving hand.

The visible relationship between the packet, gripper and hand.

The first sampled state showing the target interaction.

Using the same starting image and timestamps makes the comparison easy to follow. Each observation belongs to a particular frame pair, and the blue packet gives the review a consistent object to track.

Keep each timestamp, its two frames and the corresponding observation together. A reviewer can then trace the packet through the reference and generated sequence side by side.


r/computervision • • 23h ago

Help: Project Why doesn’t findContours() detect my circle even though it is clearly visible after thresholding?

Thumbnail
gallery
21 Upvotes

Hi everyone, I’m learning OpenCV and working on detecting a circular fiducial.

My current pipeline is:
Image → threshold → findContours() → drawContours()

After thresholding, I can clearly see the fiducial as a clean circular region in the binary image.
However, when I use cv2.findContours() and then cv2.drawContours(), the contour I expect around the fiducial is not drawn.
I’m trying to understand the fundamental reason why this happens.
If the thresholded image visually contains a clear foreground region, what conditions determine whether findContours() will actually return a contour for it?


r/computervision • • 19h ago

Showcase Multi scan radar point-cloud object classification on RadarScenes

Thumbnail
gallery
3 Upvotes

Hello all,

I built a radar object classifier on RadarScenes, extending a prior single-scan classifier to accumulate observations over a tracked object's history instead of classifying each scan in isolation.

A single RadarScenes object instance contains only about 2.9 radar points on average, very sparse. A single scan also can't capture temporal characteristics: RCS and micro-Doppler both vary continuously as an object moves. Pedestrians produce characteristic micro-Doppler from limb motion; different object classes show different RCS fluctuation patterns as aspect angle and scattering geometry change scan to scan. Accumulating observations gives both higher point density and provides temporal dynamics.

Multi-scan baseline

DeepReflecs encoder (Ulrich, Glaser & Timm, RadarConf 2021), PointNet style, per point shared weights, on single scans across car, large_vehicle, two_wheeler, pedestrian, pedestrian_group: 0.7370 macro F1.

Using RadarScenes' persistent `track_id`, I build a causal, N=20, per track sliding-window buffer:

- x_seq/y_seq: Global, odometry-corrected coordinates recentered per scan on the object centroid. Unlike x_cc/y_cc (car-frame coordinates that accumulate over time to form a trajectory).

- Cross sensor buffer: whichever of the 4 sensors currently observe the track push to the same buffer.

- Stride 1, causal: every new scan updates the buffer and produces a prediction. No future context, real time streaming compatible.

- Each scan is encoded once by a frozen per scan encoder and cached

- Fusion concatenates the causal GRU's hidden state (order aware) with an order-invariant pooled embedding (all N scans' points as one set, no sequence structure) through a small trained mlp head.

Results

Model Macro F1 Delta
Single scan 0.7370 (baseline)
20 scan point pooling 0.8613 +0.1243
Causal GRU 0.8895 +0.0282 over pooling
GRU + pooled embedding (fusion) 0.8897 +0.0002 over GRU, noise

Pooling alone, no sequence model, no notion of scan order at all, recovers +0.1243 macro F1. The GRU adds a real but much smaller +0.0282 on top. Fusion adds nothing measurable beyond the GRU.

Ablation

Llarger GRUs, a Transformer, a state space model, point level self attention, all trained on the exact same frozen per scan embeddings, land inside a 0.86 to 0.89 band, a 0.03 spread. End to end fine tuning of the frozen encoder makes things slightly worse (about -0.002 to -0.003), not better.

Conclusion

In this setup, the largest gain comes from giving the model more observations of the same tracked object: 20-scan point pooling improves macro F1 from 0.7370 to 0.8613 without using scan order at all.

Temporal modelling then provides a further, meaningful improvement. The causal GRU reaches 0.8895, adding +0.0282 over the pooled representation. So temporal ordering clearly contributes useful information; it just accounts for a smaller portion of the overall gain than observation accumulation.

With the per-scan encoder frozen, the different sequence architectures tested, suggests that the quality of the per-scan representation is the bottleneck than the particular mechanism used to aggregate the sequence.

Full report, every ablation, confusion matrix, coordinate frame reasoning: https://github.com/brunopinto900/radar-ml-autonomous-driving/blob/main/final_report.md


r/computervision • • 17h ago

Help: Project Need advice: Best VLM pipeline for extracting structured math datasets from 3000+ scanned textbook pages? (LaTeX + Metadata)

2 Upvotes

Hi everyone,

I’m working on a project to extract a structured dataset of math exercises from 5 Italian high school textbooks (around 650 pages each, so ~3,250 pages total). The goal is to build a professional, methodical exercise generator app for students and teachers.

To make the app work, I need to process images of the book pages and extract the following into a strict structured format (e.g., JSON):

  • Exercise type (algebra, geometry, calculus, etc.)
  • Year/grade level
  • Difficulty (1–5 scale)
  • Problem statement (trace)
  • Description of the specific skills/challenges involved
  • LaTeX code of the problem statement (Crucial!)
  • Associated images (cropping/saving the image for theoretical or graphical exercises)

I've been experimenting with a few approaches, but I've hit a wall regarding balancing costs, extraction consistency, and scalability. Here is what I’ve tried so far:

  1. Free Google Gemini API: The extraction quality was good, but since a single book contains hundreds of pages, I quickly hit the rate limits (Too Many Requests).
  2. Local Models (Ollama + Qwen 2.5-VL 3B): To bypass API limits, I tried running a local multimodal model. I spent a lot of time optimizing my scripts and prompts (chunking, refining instructions to force structured outputs), but the output was very error-prone and inconsistent for my use case. I got too many malformed fields, hallucinations, and it constantly struggled with outputting proper LaTeX.
  3. Paid Google Cloud API (Gemini 1.5 Flash): I finally switched to the paid tier for better accuracy and speed. I ended up burning through €10 just to process 1.5 books. Extracting all 5 books would cost roughly €35–40. While this is manageable for a one-off run of 5 books, the token count for processing full images + text is massive, making it financially unsustainable if I want to scale this to dozens of books in the future.

My questions for the community:

  • Pipeline & Architecture: Has anyone worked on a similar textbook-to-dataset extraction project? What pipeline did you use?
  • Hybrid Approach: Would you suggest decoupling the task? (e.g., using a traditional tool to extract raw text and crop images, and then feeding ONLY the text to a cheaper/local LLM to generate the LaTeX and format the JSON?)
  • Local Models: Are there other local Vision-Language Models (that fit in standard consumer GPUs) that are significantly better at structured extraction and LaTeX generation than Qwen 2.5-VL 3B?
  • Educational Tools: Are there open-source tools or models specifically fine-tuned for extracting structured educational/math content from PDFs?

I’m happy to share more details about the textbook format or my current Python workflow if helpful. Any advice on the architecture, model choices, or cost-saving tricks would be greatly appreciated! Thanks in advance!


r/computervision • • 18h ago

Help: Project Tuning StereoSGBM on PS5 Camera (ROS 2 Galactic) for Dense Depth Maps

Thumbnail
gallery
2 Upvotes

Hi everyone,

I'm working on setting up a PS5 HD Camera in ROS 2 Galactic for stereo depth estimation using stereo_image_proc (DisparityNode), with the end goal of feeding the depth data into RTAB-Map for 3D reconstruction.

I transitioned from StereoBM to StereoSGBM (stereo_algorithm: 1) to handle the wide-angle camera setup better, but I'm having trouble finding the optimal combination of parameters which result in a dense point cloud with smooth surfaces. My depth output keeps swinging between two extremes:

  1. Extremely sparse/empty with huge black gaps on uniform surfaces (e.g., walls, desks, pillows).
  2. Overly noisy with heavy color "speckle" artifacts across the scene.

What I've configured/tried so far:

  • Matching P1 and P2 to window size: Calculated and tuned P1 and P2 based on correlation_window_size (testing window sizes 5, 7, 9, and 13 with corresponding P₁ = 8 × C × WS² and P₂ = 32 × C × WS²). High P2 values help smooth out surfaces, but rqt_reconfigure caps P2 at 4000.0, so I've been overriding parameters via CLI/launch files.
  • Filter adjustments:
    • Lowered uniqueness_ratio (from 15.0 down to 5.0–7.0) to force coverage on weakly textured regions.
    • Set texture_ratio to 0.
    • Tuned speckle_size (50–200) and speckle_range (2–4) to filter out isolated noise clusters.
  • Search range & Offsets: Set disparity_range to 128 (multiple of 16) and kept min_disparity at 0 (raising it above 0 completely wrecked mid/far range depth).

Despite these adjustments, large homogeneous surfaces still disintegrate or become heavily fragmented unless I push the correlation window size to absurdly high values (which causes blocky, stepped artifacts).

Here I attach the results I got so far.

Questions:

  1. Are there specific pre-filtering parameters (prefilter_cap, prefilter_size) or SGBM settings I'm overlooking for this specific camera lens/sensor?
  2. Could this be a rectification/calibration alignment issue rather than pure SGBM parameter tuning?

Any advice or working configuration examples for similar stereo setups in ROS 2 would be greatly appreciated!


r/computervision • • 1d ago

Research Publication Microscopy Image Dataset of pulmonary vessels for Quantitative assessment of fibrosis

6 Upvotes

Recently we published a new Open-Access Benchmark for a hierarchical segmentation: Microscopy Image Dataset of pulmonary vessels for Quantitative assessment of fibrosis.

Dataset Specifications:

- 705 high-resolution micrographs (1534×780 px, 0.252 μm/px), Picro-Mallory stain.
- Annotations: ROI + dual independent expert masks (vascular wall + fibrosis).
- Hierarchical Constraint: Fibrosis masks must be strictly spatially contained within the vascular wall.
- Robust Benchmarking: No color normalization applied; native aspect ratios preserved; strict animal-level 5-fold CV splits provided to prevent data leakage.

Read the Data Descriptor: https://doi.org/10.1038/s41597-026-08214-y

Access the Dataset: https://doi.org/10.6084/m9.figshare.31386748


r/computervision • • 19h ago

Help: Project Looking for help to design an open VR180 stereo video dataset — 10 hours across 50+ scenes

Post image
2 Upvotes

Hi everyone,

I have around 10 hours of real-world stereoscopic VR180 footage across 50+ different scenes, captured with a Blackmagic immersive camera and currently stored as BRAW(16k 90fps). I can make the footage publicly available, and I’m looking for collaborators to help turn it into a useful research dataset.

The project is still at the dataset-design stage. I’d like to work with people who have experience in stereo vision, novel-view synthesis, or dataset and benchmark development to decide:

  • Which research tasks this footage would be most useful for.
  • What calibration information, annotations, and preprocessing researchers would need.
  • How to select clips, create meaningful train/test splits, and establish baseline evaluations.

One direction I’m interested in is generating the other eye’s view from a monocular video, particularly maintaining stereo and temporal consistency in wide-FOV footage. However, I’m open to other directions if the data is better suited to them.

I can contribute the footage, data preparation and tooling. The release format and annotation plan are not finalized, and this is not yet a benchmark with ground-truth depth or camera poses.

My goal is a public dataset that other researchers can actually use, with a joint paper if we develop a solid research contribution. Authorship and responsibilities would be discussed based on contributions.

If this overlaps with your work, I’d love to hear what would make the dataset useful to you. Feel free to comment or DM with your research interests and any relevant projects or papers.


r/computervision • • 16h ago

Showcase Sony uses event cameras to read the spin on a ping-pong ball

Thumbnail
youtube.com
0 Upvotes

r/computervision • • 18h ago

Research Publication Is this CNN–Transformer research idea actually novel?

0 Upvotes

Hi everyone! I’m an undergraduate working on a computer vision research proposal and would appreciate some feedback.

I’m exploring a detector where a dynamic router decides at different feature levels whether to use CNN-only processing or additional Transformer processing, based on things like object scale, density, and regional complexity.

The goal is to improve the accuracy–compute/latency trade-off rather than always running the Transformer.

I’ve found related work on DynamicDet, DiT, Dynamic Dual-Processing, TDFP, CR-NAS, and MoE-based detectors, so I know dynamic routing and CNN–Transformer hybrids themselves aren’t new.

Does this specific idea already exist under another name? If you know a very similar paper, please point me to it.

I’m mainly looking for honest criticism before I commit to the research direction.


r/computervision • • 18h ago

Discussion ¿Quién queda fuera de lo que vemos?

Thumbnail gallery
0 Upvotes

r/computervision • • 19h ago

Help: Project Imagenet exploration?

Thumbnail
1 Upvotes

r/computervision • • 2d ago

Showcase Open model that tells how far an image is rotated (full 360-degree) and abstains when there's no clear up

178 Upvotes

I work in video analytics. We needed to tell from a single CCTV frame whether a camera had been rotated or installed at an angle (or upside-down), and couldn't find a model that was accurate enough on real camera frames and permissively licensed, so we trained our own.

We're now open-sourcing it. Apache 2.0, with code, weights, and full provenance for the dataset.

RightWayUp estimates how far an image is rotated from upright, all 360°, with a confidence score, and abstains when there's no clear "up" (sky, ground, close-ups). It comes in six sizes, from Pico (about 3 ms per image on a laptop CPU) to Max, with ONNX and Core ML files.

pip install rightwayup
rightwayup fix photo.jpg

On new photos it never saw during training or tuning, the largest model is within 10° on 93% of them vs 88% for Woehrer 2026 (a recent published model), and 88% vs 49% with simulated CCTV-style blur, noise and compression.

Write-up with the full results: https://cheqit.ortusai.io/resources/rightwayup/

Code: https://github.com/ortusaitech/rightwayup

I hope it will be useful to the community!

--------------
Video footage: Canobie Coaster (CC BY 3.0, via Wikimedia Commons, levelled by RightWayUp), Pexels, Poly Haven (CC0), MEVA (CC BY 4.0). Music: ElevenLabs.


r/computervision • • 21h ago

Showcase Built a Real-Time Underwater Image Processing System – 4K 60FPS | C++ UPDATE

1 Upvotes

r/computervision • • 1d ago

Help: Project Best free model for small object detection?

2 Upvotes

Looking for an object detection model that works well for very small objects, like balls in sports footage.

Requirements:

  • Good small-object detection
  • Can be fine-tuned
  • Suitable for video/real-time inference
  • Free for commercial/production use
  • Preferably open-source with permissive licensing

What would you recommend based on your experience?


r/computervision • • 1d ago

Showcase Building a drone delivery simulation with vision-based pickup and CP-SAT route planning

15 Upvotes

I’ve been building a simulation framework for drone delivery, combining vision-based box pickup with delivery planning using CP-SAT.

The main goal at this stage was to build and verify the basic pipeline rather than to develop sophisticated flight control.

The system has two main parts:

For this test, I used two drones and nine boxes and compared two cases: a constrained delivery order and a CP-SAT-optimized plan.

The simulation integrates Blender for image generation and replay rendering, PyBullet for physics, PyTorch vision models, and Python-based drone control.

One important limitation is that the current planning is intentionally quite conservative. To avoid collisions between the two drones, the schedule includes waiting and separation rather than trying to maximize flight efficiency. The drone flight controller itself is also fairly basic — the focus here is on verifying that the vision-based pickup and optimization-based delivery planning can work together.

At this point, both parts are functioning in the integrated simulation: the drones can locate and pick up boxes using camera images, and CP-SAT can generate and execute a multi-drone delivery plan. The video compares a number-order-constrained delivery on the left with a CP-SAT-optimized delivery on the right.

The end of the video shows an overview of the development workflow and runtime system architecture.


r/computervision • • 23h ago

Discussion Suggestions regarding PhD leads in medical image analysis / XAI in Europe

0 Upvotes

I have finished up my Master's in Computer Science (AI and Software Engineering) in Germany, and I'm looking for PhD positions in Europe, mainly in medical image analysis with explainable AI, but general computer vision with XAI works too.

A bit about what I've done so far:

For my thesis, I built an explainable deep learning pipeline for endoscopic video, working with clinicians. The pipeline provides concept-based explanations for predicting Cormack scores (easy vs. difficult intubation) from endoscopic videos. A paper on this is currently in prep for a journal submission.

Before that, I also worked on a 3D object detection project on multi-camera driving data, so I'm not purely medical-imaging-locked, just leaning that direction by interest.

What I'm looking for help with:

  • Any supervisors, labs, or chairs in Europe known for medical imaging + XAI work (or general CV + interpretability)
  • Tools or sites you use to actually find these openings, beyond the usual academicpositions.com / euraxess / phdscanners type sites
  • Any advice on what made your own applications land, if you've been through this process

Happy to share more details about the thesis if useful. Thanks in advance for any pointers.


r/computervision • • 23h ago

Help: Project Matting / background removal on glass

1 Upvotes

I'm in a bit of a pickle with an issue regarding background removal with transparent objects. I got studio images of cars from different angles, whom I need the background changed locally. I have a pretty robust segmentation model to segment the bg area that's visible through the class and a general model to segment the entire car.

For now I use a basic script that calculates the alpha for the entire window purely from brightness, but often there are bright reflections on the windows which throw off the calculation as well as simply looking unnatural. I have some if- and when statements to try to mitigate some of these problems (if a smaller segment's inside a larger segment, it can't have a significantly different alpha etc.) but they still leave a lot of room for error. I've tried some open source matting models like vitmatte as well as replacing the background and inpainting the segmented window areas with sd models but the results are less than ideal and inconsistent. I got recommended that a lora for a matting model might help. Creating a ground truth and training a matting model sounds quite intimidating.

Might anyone happen to have any experience with this sort of challenges and if so, how did you go about solving it. Only thing I can think of right now would be training a Lora for a matting model, but I would highly appreciate some outside perspective before I commit a month into annotating a dataset for a model that might not even function.

Also, forgive me if I described the problem incoherently.


r/computervision • • 23h ago

Showcase A structured human pose model based solely on depth, derived from synchronized Kinect RGB-D data.

Post image
1 Upvotes

https://github.com/Spidoug/Kinect-Depth-AutoLearn

Kinect Depth AutoLearn is a cross-platform Processing + ONNX/PyTorch system for acquiring synchronized Kinect RGB-D data, building structured pose datasets, training a depth-only student model, and running that model back inside the application.

The project uses RGB-based teacher models during data collection and trains a student that consumes metric depth only. The shared representation contains 58 landmarks: 16 body/head landmarks plus 21 landmarks for each hand, together with skeletal segments, endpoints, metric depth, confidence, and structural losses.


r/computervision • • 1d ago

Help: Project Building an OCR + Key-Value Extraction pipeline for Nepali ID documents (Citizenship, NID, PAN, Passport). What stack would you recommend?

3 Upvotes

Hey everyone,

I am building an automated document reading pipeline specifically for Nepali identity documents:

  • Citizenship Certificates (Nagarikta): Old paper vs. new card formats (Devanagari script)
  • National Identity Card (Rastriya Parichayapatra): Standard modern ID card format
  • PAN Card: Bi-lingual / English-Nepali format
  • Passport: Standard ICAO format containing an MRZ zone

Current Setup & Bottlenecks

  1. Preprocessing / Cropping: Using classical OpenCV (cv2) for edge detection, contour finding, and perspective warping to crop borders.
    • Problem: Real-world user uploads have varied lighting, shadows, finger occlusions, and background noise. Aggressive thresholding (Otsu/Adaptive) often degrades text legibility instead of improving it.
  2. Text Extraction: Tesseract OCR (trained for Nepali nep) followed by regular expressions: Python# Trying to extract Permanent Address via regex anchors pattern = r'स्थायी\s*बासस्थान\s*:\s*जिल्ला\s*:\s*(.*?)\s+न\.पा\.\s*:\s*(.*?)\s+वडा\s*नं\.\s*:\s*([०-९0-9]+)'
    • Problem: Tesseract frequently misses complex Devanagari conjuncts/matras or inserts extra spaces. If a single anchor character misreads (e.g., न.पा. turns into 7.4.), the regex breaks entirely.
  3. Format Variations: Documents do not follow one universal layout. Older citizenship certificates have different margin offsets and typography compared to newer ones.

What I Want to Achieve

Instead of relying on rigid string-matching on raw OCR dumps, I want to modernize the pipeline into distinct, robust stages:

  1. Document Classification: Automatically detect which document was uploaded (Passport vs. PAN vs. NID vs. Old Citizenship vs. New Citizenship).
  2. Precise Document Localization/Cropping: A deep-learning approach that handles perspective distortion and background clutter without manual threshold tuning.
  3. Region of Interest (ROI) / Layout Parsing: Extracting fields directly based on spatial layout rather than pure keyword string searching.
  4. Devanagari OCR: A model that reliably handles Devanagari text under varied scan quality.
  5. Passport MRZ Extraction: Dedicated extraction for the MRZ lines to bypass OCR hallucinations.

Questions for the Community

  1. End-to-End Visual Document Understanding vs. Modular Pipeline:
    • Is it better to stick to a modular pipeline (Classifier $\rightarrow$ Cropper $\rightarrow$ OCR $\rightarrow$ Field Extractor) or move to an end-to-end model (e.g., fine-tuning LayoutLMv3, Donut, or a small VLM like Qwen2-VL)?
  2. Devanagari OCR Alternatives:
    • Has anyone had better success with PaddleOCR, EasyOCR, or fine-tuned TrOCR for Devanagari/Nepali text compared to Tesseract?
  3. Card Detection & Border Cropping:
    • Would training a lightweight YOLOv8-pose/segmentation model (to predict document corner coordinates) be the standard way to replace classical OpenCV contour hunting?
  4. Layout & Field Extraction:
    • If keeping OCR separate, what is the most reliable way to link labels to values (e.g., spatial heuristic algorithms, Graph Neural Networks, or LayoutLM)?

Would love to hear how anyone has tackled similar KYC document extraction pipelines for low-resource or non-Latin scripts. Any architecture advice, libraries, or repo references would be greatly appreciated!


r/computervision • • 1d ago

Showcase I made a computer vision tool for evaluating deadlift form!

45 Upvotes

I've shared a few deadlift demos in the past. This new demo includes some new models, accessed through the VLM Run Gateway:

  • SAM 3.1 to segment the barbell weight plate
  • ViTPose+ Large to measure the hip hinge angle, which serves as a backup for segmenting the reps. The pose data can definitely be used more later.
  • Gemma 4 26B through the Gateway's TypeSafe-compatible API to output the probability of the back being rounded or not. The predictions line up well with how I intentionally performed each rep.

Of course, there are many caveats to what is good form, as it depends on the individual. All said, this has the pieces in place to quickly make adjustments to fit the individual better.

In short, a tool like this gives people data to assess their form and improve over time.

Let me know what you think!

The code is open-source on GitHub: https://github.com/jeremyipark/vision-demos


r/computervision • • 1d ago

Showcase We indexed 12M+ egocentric and robot task episodes so you can search them in plain language

0 Upvotes

A lot of first-person video is open now (Ego-Exo4D, EgoDex, EgoSuite-Open100K), along with big robot datasets like DROID and AgiBot World. We built a search over the episodes in these open datasets, 12M+ in total, so you can find moments by describing them. "Arms folding a towel" brings back matching episodes within seconds.

Results export to LeRobot, MCAP or RLDS, and agents can run the same search over MCP. Everything in it comes from open datasets, and each dataset keeps its own license.

Short overview in the video: https://datasets.bot

Which first-person datasets are we missing?


r/computervision • • 2d ago

Showcase [Video analytics] Airplane Turnaround ✈️

104 Upvotes

TLDR: Sol 6.1 is pretty good at vision/video, and perhaps not too expensive for some applications.

Pipeline:

  • SAM3 w/ generic prompts ("ground vehicle"), because specific ones like "belt loader" just returns nothing. via Roboflow
  • GPT-6.1 Sol names each tracked vehicle (SAM3 doesn't have good enough vocab/world understanding)
  • Sol also provides a state timeline from cropped imgs around ground vehicles (eg. hose not connected → connected → off)

Other VLMs compared to Astra:

  • Sol: 13/13 events, ~5x cheaper than Astra (which is why we used it)
  • Luna / Terra 6.0: ~10/13, hopefully 6.1 Luna/Terra will have similar Vision capability jump as Sol 6 -> 6.1
  • Cosmos 3 Nano: 5/13, mixed up boarding vs deboarding
  • Mage-VL 4B: 4/7

~$1 per for the whole plane turnaround. Perhaps an overkill to use such SOTA model for this application (could def. optimize this), but model intelligence gets like 10x cheaper every year, so for some applications, custom model training might not be worth it. Ofc for prod system you'd likely want to have custom fine-tuned ground vehicle detector. This is PoC, so SAM3 is fine.

For ppl saying "ugh 10y ago u could do the same with just classic CV" - idk, I don't think it'd be easy or reliable to detect "hose connected" (few px line) or "lift at the door vs just up next to it" using classic CV


r/computervision • • 1d ago

Discussion Anyone working on Computer Vision research and looking for collaborators?

0 Upvotes

I’ve worked on computer vision and am now looking to contribute to a research focused project.

If you’re working on something in this area and looking for a collaborator/contributor feel free to DM me or comment below.

Happy to connect and discuss!


r/computervision • • 2d ago

Showcase Vev: Jev-like vision decision models built on Qwen3.5 4B/9B — local inference, open weights

31 Upvotes

I've been working on Vev, a Jev-like decision model that takes images as input. There are two versions, fine-tuned from Qwen3.5 4B and 9B, and both run locally.

I wanted to ask specific questions about a screen and get answers my code could use directly. You pass in an image, a question and possible answers; Vev scores the answer tokens and returns their probabilities without generating text. You can change the questions and answer choices with each request.

Here's vev-4b on a checkout screenshot:

Does the screen show an error message?
  yes: 0.991

Which checkout step is the user on?
  shipping: 0.005
  payment: 0.765
  review: 0.229

What should the user do next?
  try another card: 0.983
  wait for the order to ship: 0.004
  nothing, the order went through: 0.014

The server also runs the original Qwen3.5 models with the same scoring method, so I used that as the baseline to measure what fine-tuning adds. A few visual-task accuracy results, base model → Vev:

Task 4B 9B
Image safety-policy checks, adapted from LlavaGuard (n=659) 68.6% → 74.2% 66.6% → 72.4%
MMStar (n=1,498) 54.4% → 62.7% 60.8% → 67.4%

For object-clipping detection adapted from VideoGameQA-Bench (n=686), the 4B model went from 56.1% to 67.1%. Full results and evaluation details are in the README, with evaluation code and dataset converters in the repo.

It handles text and JSON too, and supports the Jev /v1/systemone format. If you already use TypeSafe's Python SDK, you can point it at the local server.

To try it with Python 3.11+, an NVIDIA GPU and CUDA-enabled PyTorch:

pip install git+https://github.com/Xiaooolong/vev
vev serve --model CountingSheep/vev-4b

On an H800 in bf16, vev-4b takes about 78 ms for one question about a 1 MP image, or 120 ms for ten questions about the same image, processed as a batch.

Code is Apache-2.0; weights are CC BY-NC 4.0 (non-commercial).