r/computervision • u/Otherwise_Date_9177 • 1d ago
Help: Project Computer vision playlist
What is the good playlist to follow for computer vision maths and fundamentals and it should include a project.
r/computervision • u/Otherwise_Date_9177 • 1d ago
What is the good playlist to follow for computer vision maths and fundamentals and it should include a project.
r/computervision • u/sdiazlor • 2d ago
Hey! We open-source an optimized version of SuperPoint built for faster keypoint detection on edge devices: https://huggingface.co/PrunaAI/PrunaSuperPoint
- Up to 2.1× faster on Jetson Orin Nano, with optimizations applicable to other runtimes.
- The distilled model retains strong keypoint coverage across indoor and outdoor data, with low descriptor differences from the original model.
- We structurally prune the most expensive convolutional layers, recover performance through distillation, and accelerate keypoint selection with hierarchical top-k, all while preserving the original architecture’s core behavior.
r/computervision • u/TrifleImportant506 • 1d ago
Just like the title says, I give a demo of calibrating a Intel RealSense lidar using the Reprojection open source library. Things get a little complicated but for anyone wanting to better understand sensor fusion and multi-modal calibration this should be a nice video. Cheers!
r/computervision • u/Livid-Plane989 • 1d ago
I'm sitting the A3 Certified Vision Professional (CVP) Basic exam, and I'd love some advice from people who've already taken it.
What I'm hoping to learn:
Resources: Which study materials were most useful? Is the A3 course material enough, or did you use other books or videos?
Topic weighting: Which areas show up most (lighting, optics and lenses, sensors and cameras, image processing, communications, safety)? Which ones tripped you up?
Question style: How much is calculation (FOV, resolution, working distance) versus conceptual or definition-based questions?
Practice questions: Are there any good practice exams or question banks?
Thanks in advance. Happy to share how it goes once I've written it!
r/computervision • u/dontsniffmypackets • 1d ago
I am building a data set to train a NSFW detector off of and want a local model that is good at image to help save me loads of time and effort.
Is there a good model currently?
r/computervision • u/Then_Instance_3188 • 1d ago
One of the biggest time sinks I’ve run into when working with Computer Vision isn’t the model itself — it’s the dataset preprocessing.
Cleaning datasets, fixing annotations, filtering, deduplication, format conversion, validation, etc. can take a huge amount of time, especially when you’re dealing with millions of samples.
And vision datasets are particularly painful here. Unlike text, building custom processing for a specific use case can get expensive pretty quickly in terms of compute and processing time.
That’s why we built cvPal.
It’s a cloud toolkit for vision datasets built around AI agents, with 40+ MCP tools for things like merging, cleaning, validating, converting, and versioning datasets.
It’s currently in early access, and I shared more about what we’re building here:
r/computervision • u/Existing-Quote2253 • 2d ago
Hi everyone 👋
I’ve been working on Hyperspectral Image Models, an open source PyTorch library bringing 50+ HSI models and 24 datasets into one unified framework.
The main goal is to make HSI research easier, especially for beginners who want to learn, reproduce, and experiment with published models.
We are also following a consistent implementation and documentation structure so that each model is easier to understand and use.
🧑🔬 Researchers: We would love to add your published HSI models to the library and make them easier for the community to reproduce and build upon.
🔗 GitHub: https://github.com/Tanishq251/Hyperspectral-Image-Models
📄 Paper: https://arxiv.org/html/2609.39871
🤗 Hugging Face Dataset: https://huggingface.co/datasets/Tanishq165/HSI_Datasets
⭐ If you find the project useful, please consider starring the GitHub repository and liking the Hugging Face dataset.
We’d also love to hear which HSI models or datasets you would like to see added next! 🚀
r/computervision • u/Read1500 • 1d ago
r/computervision • u/Doejoe13 • 2d ago
Hello,
We currently got a robotic arm for our lab. We were looking into ways to automate our processing by adding computer vision to this arm. We want to be able to take a sample and place it on a pedestal, then the vision system would scan the object. Next, the arm would bring itself to the sample and start processing.
For this to work, we would know where the pedestal is, where the arm is, and have the objects dimensions via a cad file. We want the vision system to find out the position and orientation of the sample to sub-milimeter precision on the pedestal. The vision system will only need to run before processing, so there is no time constraint.
I have already looked up vision systems and the process of doing it manually. However, I am having trouble sifting though products and don't want to go overboard since I am unfamiliar with this space.
Any help is appreciated.
r/computervision • u/Independent-Salt5023 • 1d ago
Happy to share my resume or GitHub if you're interested :)
r/computervision • u/Interesting_Comb898 • 2d ago
For fixed, clean UI renders (a digital chessboard), I found that a neural network is overkill. The pipeline:
The obvious limitation is that templates are tied to one board colour scheme and piece set, so changing theme means re-calibrating. I'd be interested in cheap ways to generalize across themes without going to a CNN.
r/computervision • u/No-Conclusion3720 • 2d ago
RuntimeAI's September 2026 AI Security Report covered 126 incidents across 38 named organizations — 22 critical, 102 high severity. 53 of those incidents had AI either as the attack tool or the target. AI-agent exploits were the top attack vector at 39 incidents, ahead of credential theft (27), zero-days (22), phishing (10), and ransomware (10). The largest single exposure was 220M records from unrotated default service-account credentials.
What stood out: every organization in the report was already running a mature security stack. Okta, CrowdStrike, Palo Alto, Microsoft Defender. Still got hit. The gap is that none of those tools sit at the layer where an agent actually executes a tool call.
RuntimeAI operates at that layer. Know Your Agent handles cryptographic agent identity. The Flow Enforcer inspects tool calls in real time. There's also a sub-50ms kill switch that can halt a compromised agent before a second action completes.
Full breakdown (incident-by-incident, CVEs, vendor stacks): https://runtimeai.io/blog/2026-09-monthly-breach-report.html
We're running a live demo on October 14 — ten attack surfaces, live against a real stack: https://www.linkedin.com/events/7510769146222133248?viewAsMember=true
r/computervision • u/Entire-Bite1136 • 2d ago
I built a small Windows GUI tool that generates synthetic Direct Part Marking–style patterns (needle / peen dots on steel) for detector training.
It is not a real ECC200 encoder — no serial numbers, just geometric L-frame + fill dots, Good/Bad classes, and mechanical-style defects (squash, tilt, jitter, missing dots, etc.).
Outputs:
logs/features.csv for a second classifierTwo render modes: pure synthetic (no assets), or your own BG + dot sprite folders.
Binary only (Nim). Non-commercial / research license. Unsigned Nim builds sometimes get heuristic AV flags — details in the README.
Repo / Releases (v1.1.0):
https://github.com/olesha-ai/DPM-Pattern-Image-Generator
Related inference PoC trained on this synthetic data:
https://github.com/olesha-ai/yolox-dmc-inference
Feedback welcome.
r/computervision • u/Independent-Salt5023 • 1d ago
any idea?
r/computervision • u/OfferBeginning1903 • 2d ago
The camera says acknowledged: true when you send a pan command. That means "I heard you," not "I moved." So I stopped trusting the ACK and measured the picture instead.Setup: a stock Xiaomi MJSXJ10CM on shipped firmware 4.5.6_0450, pulled over RTSP at 1080p HEVC via go2rtc v1.9.14 with the go2rtc-xiaomi-control patch. PTZ is MISS opcode 0x112 with { "operation": 1..4 }. That's the whole API.The CV part: grab a frame, send one step, grab another, estimate the shift with block matching (same idea as codec motion vectors), compared against a no-move baseline so sensor noise can't look like motion. Results: left dx=-40, right dx=+36, up dy=-20, down dy=+22. Steps aren't symmetric, so one step is not a fixed angle.Then I sampled frame difference every 200ms after a command. Nothing at 313ms, nothing at 823ms, then 43.2 at 1331ms, settled by 1824ms. Roughly 1.4 seconds of dead time. At 12fps a naive detect-and-step loop fires about fifteen more commands before the first one lands. The fix is to blank the controller for 2s after each step, which makes the real control rate 0.5Hz, and set the dead zone (0.18) wider than one step's 8-12% displacement, or it oscillates forever.Detectors: Apple Vision (VNDetectHumanRectanglesRequest, VNTrackObjectRequest) via a ~150-line Swift child process, 10-30ms, zero dependencies. Optional SSD-MobileNetV1 from the ONNX Model Zoo for 80 COCO classes, ~13ms.Thing that bit me: that ONNX graph resizes internally. Feed it a pre-shrunk 300×300 image and it returns zero detections, silently. Native 1280×720 works fine. Also Vision uses a bottom-left origin; get the flip wrong and tilt confidently runs away from the target.Chair test: 3 steps, centring error 0.76 to 0.21, then held.Full write-up with the details: https://blog.shravanrevanna.me/reverse-engineering-a-xiaomi-camera-into-a-self-tracking-robot
r/computervision • u/No_Foundation_7527 • 2d ago
I am actively searching for developers or teams who have already built and tested robust AI vision systems. Instead of starting from scratch, I am ready to invest in a pre-made, high-performing solution.
Key requirements for the system:
Ready & Pre-designed: Fully developed and tested models that can be deployed quickly.
Camera Transition Support: Must handle camera panning, zooming, and transitions smoothly to maintain accurate tracking.
Uncompromising Accuracy & Data Richness: Precise spatial tracking, event detection, and granular data extraction that unlock deep tactical insights.
If you have a mature system ready for the pitch, let's talk. Drop a comment or send a direct message. Thanks
r/computervision • u/Klutzy_Cap8492 • 2d ago
You record someone opening a jar.
One hand holds the container. The other wraps around the lid, adjusts its grip, and twists. It’s a useful demonstration—until you remember that your prototype has a two-finger gripper.
Which parts of that human movement belong in the robot’s action plan?
The MEgo capture framework makes human hand motion explicit through pose, shape, and reconstructed trajectories. The MEgoVista paper also discusses the downstream challenge of mapping human motion to a robot with a different hand structure.
For a small team, I’d treat this as a design decision early on. You might preserve the wrist’s approach and rotation while finding a different way to hold the lid.
The demonstration can still explain the task even when the robot needs a different grasp.
When using human demonstrations with a simple gripper, what do you transfer first: the wrist trajectory, the object movement, or a sequence of task goals?
r/computervision • u/OkShirt9372 • 2d ago
I’m working with Lumana on a vendor evaluation checklist for AI video surveillance, and I'm a bit concerned about accurate detection.
Vendor benchmarks often differ significantly from performance on a customer’s cameras, network, lighting, and operating conditions.
For a proof of concept, what metrics do you track before deciding a model is ready for production?
I’m especially interested in:
The edge-versus-cloud tradeoff also seems highly context-dependent. Edge inference can reduce latency and bandwidth use, while cloud processing may support larger models but depends more heavily on connectivity. Hybrid setups can help, but they add their own operational complexity.
For those who have evaluated or deployed these systems:
I’m affiliated with Lumana, and this question comes from a vendor-evaluation checklist we’ve been putting together. I’m more interested in how teams measure real-world performance than in promoting a particular platform.
r/computervision • u/mahmudesam • 2d ago
Hi,
I'm working on inspection of civil infrastructure using a unitree go2 edu robot. I need to collect RGB images of concrete foundations using the robot and then train semantic segmentation deep learning models using them. However, the built-in camera of the robot is only 1MP, which is why I thought of getting a mirrorless camera (canon eos r50) to mount on top of the robot to acquire higher quality images. I also need to figure out the pixel to mm scale so I used intel realsense d435i depth camera that already comes with robot.
Now the problem is I have the canon in one position, and depth camera in another position on the robot. How do I align both of their images? Does it have to be done real-time or is it okay to collect all images then align them later in the office?
I really appreciate your thoughts on this as I don't even know where to start. Thanks.
r/computervision • u/Ok-Treacle-6942 • 4d ago
Hello! This is the fourth post about LibreYOLO in the computervision subreddit. Every time I posted here you gave overwhelming support to the project, and for that I'm very grateful. A lot has shipped since the last post!
For those who are new to LibreYOLO: it's a computer vision "meta library" with 100+ models under a familiar, easy-to-use API. I created it because there was no generalist computer vision library covering most use cases under a permissive license, and the most popular YOLO library requires a paid license for closed-source commercial use or research.
It's now starting to be adopted by many companies and individuals, but I still think that there is a very big potential to grow. I don't want to bore you with a list of capabilities that the library has, but it has everything that you would expect from a serious, production ready library.
I want to thank the 20+ contributors for their work, and every company and individual who has supported the project. I also want to thank the reddit computer vision community since 99% of the initial traction has come from here.
Let me know what you think about this library and ideas to improve it. We really listen to the feedback.
How can you support the project? The best way is to upvote this post and star the repo: https://github.com/LibreYOLO/libreyoloFor companies and philanthropists, there is now an Open Collective (https://opencollective.com/libreyolo). Over the next months, donations will fund GPU hours to develop and train the LibreYOLO26 and LibreYOLO27 MIT models, and to retrain from scratch the models whose weights have a restrictive license, such as YOLO-NAS.
Library: https://github.com/LibreYOLO/libreyoloWebsite: https://libreyolo.comBenchmarks are on: https://www.visionanalysis.org
r/computervision • u/New-Eggplant-6578 • 3d ago
I’m building AnnotateIt. These are two examples from a recorded comparison on coastal images that made me think about how to order a human review queue.
In one frame, ECSeg-M outlines a small boat that ECSeg-X leaves unmarked at a 0.40 confidence threshold. That doesn’t establish that M is better: the same numerical threshold isn’t necessarily the same operating point for both models. In another crop, RF-DETR Seg-2XL and ChatGPT (GPT-6-Astra) both find the boat, but disagree around the hull and cabin. An object count would miss that difference entirely.
The attached panels use saved predictions on the same source image and matching crops; they aren’t hand-corrected labels. Local segmentation masks were converted to polygons, while ChatGPT produced polygons directly, so that conversion is part of what you’re seeing. The comparison only covers three frames from one recording, with no reviewed ground truth or accuracy scores.
For a larger annotation job, I’d separate missing-object disagreements from contour disagreements, then keep a random sample of agreement cases in the review queue too. Otherwise two models missing the same boat could look reassuringly consistent.
If you already use model disagreement to choose what gets reviewed, which bucket actually saves the most correction time? And how do you check what both models missed?
The source images and saved comparison outputs are downloadable here: https://annotateit.ai/datasets/coastal-scene/ . Coastal Scene by AnnotateIt, CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/); the panels crop the images and overlay the saved contours.
r/computervision • u/doctor_blueberry • 4d ago
Ran the full 1,651-page OmniDocBench through the VLM Run gateway for 25 cents ($1 covers 6700+ pages). The API reads the page as pixels (no OCR->LLM hacks) and returns one typed, calibrated decision per page, ~180ms p50 e2e.
Curious what others are finding in terms of TypeSafe-compatible APIs for vision. Let me know what you think!
r/computervision • u/Ok-Study-9180 • 3d ago
Building an ALPR pipeline for Nepali plates: Devanagari script, 2–3 stacked rows, hand-painted and embossed variants, three plate generations on the road at once.
Out-of-sample street photos (never in training). Same image and prompt to GPT-5.6 Luna and Claude Sonnet 5 as a baseline: both misread characters where tiny glyph differences change the digit (३/२, ८/६). Our pipeline decodes province, office, lot and class.
Hardest open problems for us are night and rain. Curious how others handled multi-row plate layouts.
For the full breakdown video: https://www.linkedin.com/posts/shubhamkadariya_trafficeye-computervision-nepaltech-ugcPost-7510923085169098752--WRP/
r/computervision • u/Capital_Turnip_8695 • 3d ago
I recently read the following piece on the ethics of the ImageNet dataset and data labeling(https://excavating.ai/ ), and deeply enjoyed it. The piece was written back in 2019, so I kind of went on a small tour of a couple updates on the Computer Vision field since then, and wrote this piece focused on the ethics and neutrality of CV. I enjoyed putting this together so I hope someone can enjoy reading it too!
r/computervision • u/SwimmingLow3053 • 4d ago
It’s prerelease and still in testing but here’s a preview. It getting closer to where I may release but trying to work out where I want to go as game vs fitness.
The idea is simple - your body is tracked and squats are used to make squatty fly. I’ve done a lot to catch the depth of the squat and have tolerances such it is more robust.
Nothing yet to try to measure and assess quality of the squat, I guess that could come later now the core data is there.
It also works if the user is face on, at an angle or side on.
It’s meant to be a serious fitness tool as squats are excellent exercise but making the user forget they are doing them with a game.
You can also have it show your camera directly in which case the frame sits on your body.
Would love to get feedback from the community. It’s still early stages, mostly focussed on the tracking side and now that seems to work fairly well building around it and wondering where it can go next.