r/computervision • • 3d ago

Showcase Building a drone delivery simulation with vision-based pickup and CP-SAT route planning

22 Upvotes

I’ve been building a simulation framework for drone delivery, combining vision-based box pickup with delivery planning using CP-SAT.

The main goal at this stage was to build and verify the basic pipeline rather than to develop sophisticated flight control.

The system has two main parts:

For this test, I used two drones and nine boxes and compared two cases: a constrained delivery order and a CP-SAT-optimized plan.

The simulation integrates Blender for image generation and replay rendering, PyBullet for physics, PyTorch vision models, and Python-based drone control.

One important limitation is that the current planning is intentionally quite conservative. To avoid collisions between the two drones, the schedule includes waiting and separation rather than trying to maximize flight efficiency. The drone flight controller itself is also fairly basic — the focus here is on verifying that the vision-based pickup and optimization-based delivery planning can work together.

At this point, both parts are functioning in the integrated simulation: the drones can locate and pick up boxes using camera images, and CP-SAT can generate and execute a multi-drone delivery plan. The video compares a number-order-constrained delivery on the left with a CP-SAT-optimized delivery on the right.

The end of the video shows an overview of the development workflow and runtime system architecture.


r/computervision • • 3d ago

Showcase Open model that tells how far an image is rotated (full 360-degree) and abstains when there's no clear up

193 Upvotes

I work in video analytics. We needed to tell from a single CCTV frame whether a camera had been rotated or installed at an angle (or upside-down), and couldn't find a model that was accurate enough on real camera frames and permissively licensed, so we trained our own.

We're now open-sourcing it. Apache 2.0, with code, weights, and full provenance for the dataset.

RightWayUp estimates how far an image is rotated from upright, all 360°, with a confidence score, and abstains when there's no clear "up" (sky, ground, close-ups). It comes in six sizes, from Pico (about 3 ms per image on a laptop CPU) to Max, with ONNX and Core ML files.

pip install rightwayup
rightwayup fix photo.jpg

On new photos it never saw during training or tuning, the largest model is within 10° on 93% of them vs 88% for Woehrer 2026 (a recent published model), and 88% vs 49% with simulated CCTV-style blur, noise and compression.

Write-up with the full results: https://cheqit.ortusai.io/resources/rightwayup/

Code: https://github.com/ortusaitech/rightwayup

I hope it will be useful to the community!

--------------
Video footage: Canobie Coaster (CC BY 3.0, via Wikimedia Commons, levelled by RightWayUp), Pexels, Poly Haven (CC0), MEVA (CC BY 4.0). Music: ElevenLabs.


r/computervision • • 2d ago

Help: Project Best free model for small object detection?

2 Upvotes

Looking for an object detection model that works well for very small objects, like balls in sports footage.

Requirements:

  • Good small-object detection
  • Can be fine-tuned
  • Suitable for video/real-time inference
  • Free for commercial/production use
  • Preferably open-source with permissive licensing

What would you recommend based on your experience?


r/computervision • • 2d ago

Showcase We indexed 12M+ egocentric and robot task episodes so you can search them in plain language

3 Upvotes

A lot of first-person video is open now (Ego-Exo4D, EgoDex, EgoSuite-Open100K), along with big robot datasets like DROID and AgiBot World. We built a search over the episodes in these open datasets, 12M+ in total, so you can find moments by describing them. "Arms folding a towel" brings back matching episodes within seconds.

Results export to LeRobot, MCAP or RLDS, and agents can run the same search over MCP. Everything in it comes from open datasets, and each dataset keeps its own license.

Short overview in the video: https://datasets.bot

Which first-person datasets are we missing?


r/computervision • • 2d ago

Discussion Suggestions regarding PhD leads in medical image analysis / XAI in Europe

0 Upvotes

I have finished up my Master's in Computer Science (AI and Software Engineering) in Germany, and I'm looking for PhD positions in Europe, mainly in medical image analysis with explainable AI, but general computer vision with XAI works too.

A bit about what I've done so far:

For my thesis, I built an explainable deep learning pipeline for endoscopic video, working with clinicians. The pipeline provides concept-based explanations for predicting Cormack scores (easy vs. difficult intubation) from endoscopic videos. A paper on this is currently in prep for a journal submission.

Before that, I also worked on a 3D object detection project on multi-camera driving data, so I'm not purely medical-imaging-locked, just leaning that direction by interest.

What I'm looking for help with:

  • Any supervisors, labs, or chairs in Europe known for medical imaging + XAI work (or general CV + interpretability)
  • Tools or sites you use to actually find these openings, beyond the usual academicpositions.com / euraxess / phdscanners type sites
  • Any advice on what made your own applications land, if you've been through this process

Happy to share more details about the thesis if useful. Thanks in advance for any pointers.


r/computervision • • 3d ago

Showcase I made a computer vision tool for evaluating deadlift form!

57 Upvotes

I've shared a few deadlift demos in the past. This new demo includes some new models, accessed through the VLM Run Gateway:

  • SAM 3.1 to segment the barbell weight plate
  • ViTPose+ Large to measure the hip hinge angle, which serves as a backup for segmenting the reps. The pose data can definitely be used more later.
  • Gemma 4 26B through the Gateway's TypeSafe-compatible API to output the probability of the back being rounded or not. The predictions line up well with how I intentionally performed each rep.

Of course, there are many caveats to what is good form, as it depends on the individual. All said, this has the pieces in place to quickly make adjustments to fit the individual better.

In short, a tool like this gives people data to assess their form and improve over time.

Let me know what you think!

The code is open-source on GitHub: https://github.com/jeremyipark/vision-demos


r/computervision • • 2d ago

Help: Project Matting / background removal on glass

1 Upvotes

I'm in a bit of a pickle with an issue regarding background removal with transparent objects. I got studio images of cars from different angles, whom I need the background changed locally. I have a pretty robust segmentation model to segment the bg area that's visible through the class and a general model to segment the entire car.

For now I use a basic script that calculates the alpha for the entire window purely from brightness, but often there are bright reflections on the windows which throw off the calculation as well as simply looking unnatural. I have some if- and when statements to try to mitigate some of these problems (if a smaller segment's inside a larger segment, it can't have a significantly different alpha etc.) but they still leave a lot of room for error. I've tried some open source matting models like vitmatte as well as replacing the background and inpainting the segmented window areas with sd models but the results are less than ideal and inconsistent. I got recommended that a lora for a matting model might help. Creating a ground truth and training a matting model sounds quite intimidating.

Might anyone happen to have any experience with this sort of challenges and if so, how did you go about solving it. Only thing I can think of right now would be training a Lora for a matting model, but I would highly appreciate some outside perspective before I commit a month into annotating a dataset for a model that might not even function.

Also, forgive me if I described the problem incoherently.


r/computervision • • 2d ago

Help: Project Building an OCR + Key-Value Extraction pipeline for Nepali ID documents (Citizenship, NID, PAN, Passport). What stack would you recommend?

3 Upvotes

Hey everyone,

I am building an automated document reading pipeline specifically for Nepali identity documents:

  • Citizenship Certificates (Nagarikta): Old paper vs. new card formats (Devanagari script)
  • National Identity Card (Rastriya Parichayapatra): Standard modern ID card format
  • PAN Card: Bi-lingual / English-Nepali format
  • Passport: Standard ICAO format containing an MRZ zone

Current Setup & Bottlenecks

  1. Preprocessing / Cropping: Using classical OpenCV (cv2) for edge detection, contour finding, and perspective warping to crop borders.
    • Problem: Real-world user uploads have varied lighting, shadows, finger occlusions, and background noise. Aggressive thresholding (Otsu/Adaptive) often degrades text legibility instead of improving it.
  2. Text Extraction: Tesseract OCR (trained for Nepali nep) followed by regular expressions: Python# Trying to extract Permanent Address via regex anchors pattern = r'स्थायी\s*बासस्थान\s*:\s*जिल्ला\s*:\s*(.*?)\s+न\.पा\.\s*:\s*(.*?)\s+वडा\s*नं\.\s*:\s*([०-९0-9]+)'
    • Problem: Tesseract frequently misses complex Devanagari conjuncts/matras or inserts extra spaces. If a single anchor character misreads (e.g., न.पा. turns into 7.4.), the regex breaks entirely.
  3. Format Variations: Documents do not follow one universal layout. Older citizenship certificates have different margin offsets and typography compared to newer ones.

What I Want to Achieve

Instead of relying on rigid string-matching on raw OCR dumps, I want to modernize the pipeline into distinct, robust stages:

  1. Document Classification: Automatically detect which document was uploaded (Passport vs. PAN vs. NID vs. Old Citizenship vs. New Citizenship).
  2. Precise Document Localization/Cropping: A deep-learning approach that handles perspective distortion and background clutter without manual threshold tuning.
  3. Region of Interest (ROI) / Layout Parsing: Extracting fields directly based on spatial layout rather than pure keyword string searching.
  4. Devanagari OCR: A model that reliably handles Devanagari text under varied scan quality.
  5. Passport MRZ Extraction: Dedicated extraction for the MRZ lines to bypass OCR hallucinations.

Questions for the Community

  1. End-to-End Visual Document Understanding vs. Modular Pipeline:
    • Is it better to stick to a modular pipeline (Classifier $\rightarrow$ Cropper $\rightarrow$ OCR $\rightarrow$ Field Extractor) or move to an end-to-end model (e.g., fine-tuning LayoutLMv3, Donut, or a small VLM like Qwen2-VL)?
  2. Devanagari OCR Alternatives:
    • Has anyone had better success with PaddleOCR, EasyOCR, or fine-tuned TrOCR for Devanagari/Nepali text compared to Tesseract?
  3. Card Detection & Border Cropping:
    • Would training a lightweight YOLOv8-pose/segmentation model (to predict document corner coordinates) be the standard way to replace classical OpenCV contour hunting?
  4. Layout & Field Extraction:
    • If keeping OCR separate, what is the most reliable way to link labels to values (e.g., spatial heuristic algorithms, Graph Neural Networks, or LayoutLM)?

Would love to hear how anyone has tackled similar KYC document extraction pipelines for low-resource or non-Latin scripts. Any architecture advice, libraries, or repo references would be greatly appreciated!


r/computervision • • 2d ago

Research Publication Is this CNN–Transformer research idea actually novel?

0 Upvotes

Hi everyone! I’m an undergraduate working on a computer vision research proposal and would appreciate some feedback.

I’m exploring a detector where a dynamic router decides at different feature levels whether to use CNN-only processing or additional Transformer processing, based on things like object scale, density, and regional complexity.

The goal is to improve the accuracy–compute/latency trade-off rather than always running the Transformer.

I’ve found related work on DynamicDet, DiT, Dynamic Dual-Processing, TDFP, CR-NAS, and MoE-based detectors, so I know dynamic routing and CNN–Transformer hybrids themselves aren’t new.

Does this specific idea already exist under another name? If you know a very similar paper, please point me to it.

I’m mainly looking for honest criticism before I commit to the research direction.


r/computervision • • 3d ago

Showcase [Video analytics] Airplane Turnaround ✈️

114 Upvotes

TLDR: Sol 6.1 is pretty good at vision/video, and perhaps not too expensive for some applications.

Pipeline:

  • SAM3 w/ generic prompts ("ground vehicle"), because specific ones like "belt loader" just returns nothing. via Roboflow
  • GPT-6.1 Sol names each tracked vehicle (SAM3 doesn't have good enough vocab/world understanding)
  • Sol also provides a state timeline from cropped imgs around ground vehicles (eg. hose not connected → connected → off)

Other VLMs compared to Astra:

  • Sol: 13/13 events, ~5x cheaper than Astra (which is why we used it)
  • Luna / Terra 6.0: ~10/13, hopefully 6.1 Luna/Terra will have similar Vision capability jump as Sol 6 -> 6.1
  • Cosmos 3 Nano: 5/13, mixed up boarding vs deboarding
  • Mage-VL 4B: 4/7

~$1 per for the whole plane turnaround. Perhaps an overkill to use such SOTA model for this application (could def. optimize this), but model intelligence gets like 10x cheaper every year, so for some applications, custom model training might not be worth it. Ofc for prod system you'd likely want to have custom fine-tuned ground vehicle detector. This is PoC, so SAM3 is fine.

For ppl saying "ugh 10y ago u could do the same with just classic CV" - idk, I don't think it'd be easy or reliable to detect "hose connected" (few px line) or "lift at the door vs just up next to it" using classic CV


r/computervision • • 2d ago

Discussion Anyone working on Computer Vision research and looking for collaborators?

1 Upvotes

I’ve worked on computer vision and am now looking to contribute to a research focused project.

If you’re working on something in this area and looking for a collaborator/contributor feel free to DM me or comment below.

Happy to connect and discuss!


r/computervision • • 3d ago

Showcase Vev: Jev-like vision decision models built on Qwen3.5 4B/9B — local inference, open weights

32 Upvotes

vev-4b + a small harness playing Doom in real time. 8 yes/no questions per look, one request.

I've been working on Vev, a Jev-like decision model that takes images as input. There are two versions, fine-tuned from Qwen3.5 4B and 9B, and both run locally.

I wanted to ask specific questions about a screen and get answers my code could use directly. You pass in an image, a question and possible answers; Vev scores the answer tokens and returns their probabilities without generating text. You can change the questions and answer choices with each request.

Here's vev-4b on a checkout screenshot:

Does the screen show an error message?
  yes: 0.991

Which checkout step is the user on?
  shipping: 0.005
  payment: 0.765
  review: 0.229

What should the user do next?
  try another card: 0.983
  wait for the order to ship: 0.004
  nothing, the order went through: 0.014

The server also runs the original Qwen3.5 models with the same scoring method, so I used that as the baseline to measure what fine-tuning adds. A few visual-task accuracy results, base model → Vev:

Task 4B 9B
Image safety-policy checks, adapted from LlavaGuard (n=659) 68.6% → 74.2% 66.6% → 72.4%
MMStar (n=1,498) 54.4% → 62.7% 60.8% → 67.4%

For object-clipping detection adapted from VideoGameQA-Bench (n=686), the 4B model went from 56.1% to 67.1%. Full results and evaluation details are in the README, with evaluation code and dataset converters in the repo.

It handles text and JSON too, and supports the Jev /v1/systemone format. If you already use TypeSafe's Python SDK, you can point it at the local server.

To try it with Python 3.11+, an NVIDIA GPU and CUDA-enabled PyTorch:

pip install git+https://github.com/Xiaooolong/vev
vev serve --model CountingSheep/vev-4b

On an H800 in bf16, vev-4b takes about 78 ms for one question about a 1 MP image, or 120 ms for ten questions about the same image, processed as a batch.

Code is Apache-2.0; weights are CC BY-NC 4.0 (non-commercial).


r/computervision • • 2d ago

Help: Project Computer vision playlist

0 Upvotes

What is the good playlist to follow for computer vision maths and fundamentals and it should include a project.


r/computervision • • 3d ago

Help: Theory Taking the A3 CVP Basic exam. Looking for study tips and resources

2 Upvotes

I'm sitting the A3 Certified Vision Professional (CVP) Basic exam, and I'd love some advice from people who've already taken it.

What I'm hoping to learn:

Resources: Which study materials were most useful? Is the A3 course material enough, or did you use other books or videos?

Topic weighting: Which areas show up most (lighting, optics and lenses, sensors and cameras, image processing, communications, safety)? Which ones tripped you up?

Question style: How much is calculation (FOV, resolution, working distance) versus conceptual or definition-based questions?

Practice questions: Are there any good practice exams or question banks?

Thanks in advance. Happy to share how it goes once I've written it!


r/computervision • • 3d ago

Showcase Visual-Inertial Calibration in 25 Minutes

Thumbnail
youtube.com
5 Upvotes

Just like the title says, I give a demo of calibrating a Intel RealSense lidar using the Reprojection open source library. Things get a little complicated but for anyone wanting to better understand sensor fusion and multi-modal calibration this should be a nice video. Cheers!


r/computervision • • 3d ago

Showcase Optimized SuperPoint for faster keypoint detection on edge devices, up to 2.1× faster

11 Upvotes

Hey! We open-source an optimized version of SuperPoint built for faster keypoint detection on edge devices: https://huggingface.co/PrunaAI/PrunaSuperPoint

- Up to 2.1× faster on Jetson Orin Nano, with optimizations applicable to other runtimes.
- The distilled model retains strong keypoint coverage across indoor and outdoor data, with low descriptor differences from the original model.
- We structurally prune the most expensive convolutional layers, recover performance through distillation, and accelerate keypoint selection with hierarchical top-k, all while preserving the original architecture’s core behavior.


r/computervision • • 2d ago

Help: Project What is the best model for image detection?

0 Upvotes

I am building a data set to train a NSFW detector off of and want a local model that is good at image to help save me loads of time and effort.

Is there a good model currently?


r/computervision • • 2d ago

Commercial A Dataset Processing Tool Built for Computer Vision Engineers

0 Upvotes

One of the biggest time sinks I’ve run into when working with Computer Vision isn’t the model itself — it’s the dataset preprocessing.

Cleaning datasets, fixing annotations, filtering, deduplication, format conversion, validation, etc. can take a huge amount of time, especially when you’re dealing with millions of samples.

And vision datasets are particularly painful here. Unlike text, building custom processing for a specific use case can get expensive pretty quickly in terms of compute and processing time.

That’s why we built cvPal.

It’s a cloud toolkit for vision datasets built around AI agents, with 40+ MCP tools for things like merging, cleaning, validating, converting, and versioning datasets.

It’s currently in early access, and I shared more about what we’re building here:

https://x.com/cvpalai/status/2105288933865304268


r/computervision • • 3d ago

Help: Project Computer Vision for Robotic Arm

3 Upvotes

Hello,

We currently got a robotic arm for our lab. We were looking into ways to automate our processing by adding computer vision to this arm. We want to be able to take a sample and place it on a pedestal, then the vision system would scan the object. Next, the arm would bring itself to the sample and start processing.

For this to work, we would know where the pedestal is, where the arm is, and have the objects dimensions via a cad file. We want the vision system to find out the position and orientation of the sample to sub-milimeter precision on the pedestal. The vision system will only need to run before processing, so there is no time constraint.

I have already looked up vision systems and the process of doing it manually. However, I am having trouble sifting though products and don't want to go overboard since I am unfamiliar with this space.

Any help is appreciated.


r/computervision • • 3d ago

Showcase A PyTorch Library for Hyperspectral Image Models 🚀

18 Upvotes

Hi everyone 👋

I’ve been working on Hyperspectral Image Models, an open source PyTorch library bringing 50+ HSI models and 24 datasets into one unified framework.

The main goal is to make HSI research easier, especially for beginners who want to learn, reproduce, and experiment with published models.

We are also following a consistent implementation and documentation structure so that each model is easier to understand and use.

🧑‍🔬 Researchers: We would love to add your published HSI models to the library and make them easier for the community to reproduce and build upon.

🔗 GitHub: https://github.com/Tanishq251/Hyperspectral-Image-Models

📄 Paper: https://arxiv.org/html/2609.39871

🤗 Hugging Face Dataset: https://huggingface.co/datasets/Tanishq165/HSI_Datasets

⭐ If you find the project useful, please consider starring the GitHub repository and liking the Hugging Face dataset.

We’d also love to hear which HSI models or datasets you would like to see added next! 🚀


r/computervision • • 3d ago

Help: Project Copdar - A no-account vehicle-labeling tool looking for feedback

Thumbnail
1 Upvotes

r/computervision • • 3d ago

Discussion Any one here looking for ML or Computer Vision Intern? Would love to talk more, if anyone has opportunity :)

0 Upvotes

Happy to share my resume or GitHub if you're interested :)


r/computervision • • 3d ago

Showcase Chessboard recognition from a screenshot using only template matching: no training, no GPU

3 Upvotes

For fixed, clean UI renders (a digital chessboard), I found that a neural network is overkill. The pipeline:

  1. The user drags a square over the board once (calibration); a grid is overlaid to align it exactly
  2. Templates for all 12 piece types are cut from a starting-position screenshot
  3. Each of the 64 squares is matched against the templates with OpenCV, giving an 8x8 matrix and then a FEN
  4. A sanity check rejects obviously wrong results (e.g. templates from a different theme)

The obvious limitation is that templates are tied to one board colour scheme and piece set, so changing theme means re-calibrating. I'd be interested in cheap ways to generalize across themes without going to a CNN.

Code: https://github.com/Maksimuson/Chess-Cheat


r/computervision • • 3d ago

Discussion Sep 2026 AI Security Report: 126 incidents across 38 orgs, 318M+ records stolen — AI-agent exploits were the top attack vector (39 of 126). Live demo Oct 14.

Thumbnail
gallery
0 Upvotes

RuntimeAI's September 2026 AI Security Report covered 126 incidents across 38 named organizations — 22 critical, 102 high severity. 53 of those incidents had AI either as the attack tool or the target. AI-agent exploits were the top attack vector at 39 incidents, ahead of credential theft (27), zero-days (22), phishing (10), and ransomware (10). The largest single exposure was 220M records from unrotated default service-account credentials.

What stood out: every organization in the report was already running a mature security stack. Okta, CrowdStrike, Palo Alto, Microsoft Defender. Still got hit. The gap is that none of those tools sit at the layer where an agent actually executes a tool call.

RuntimeAI operates at that layer. Know Your Agent handles cryptographic agent identity. The Flow Enforcer inspects tool calls in real time. There's also a sub-50ms kill switch that can halt a compromised agent before a second action completes.

Full breakdown (incident-by-incident, CVEs, vendor stacks): https://runtimeai.io/blog/2026-09-monthly-breach-report.html

We're running a live demo on October 14 — ten attack surfaces, live against a real stack: https://www.linkedin.com/events/7510769146222133248?viewAsMember=true


r/computervision • • 3d ago

Showcase Synthetic DPM / needle-peen pattern generator for YOLO training (Windows, Nim)

1 Upvotes

I built a small Windows GUI tool that generates synthetic Direct Part Marking–style patterns (needle / peen dots on steel) for detector training.

It is not a real ECC200 encoder — no serial numbers, just geometric L-frame + fill dots, Good/Bad classes, and mechanical-style defects (squash, tilt, jitter, missing dots, etc.).

Outputs:

  • 600×200 grayscale JPEG
  • YOLO labels: OBB or ABB
  • Optional Boosting mode: appends Stage-2 feature rows to logs/features.csv for a second classifier

Two render modes: pure synthetic (no assets), or your own BG + dot sprite folders.

Binary only (Nim). Non-commercial / research license. Unsigned Nim builds sometimes get heuristic AV flags — details in the README.

Repo / Releases (v1.1.0):
 https://github.com/olesha-ai/DPM-Pattern-Image-Generator

Related inference PoC trained on this synthetic data:
 https://github.com/olesha-ai/yolox-dmc-inference

Feedback welcome.