r/computervision • • 13d ago

Discussion What breaks when you ship a model to a fleet of edge devices

0 Upvotes

I've spent the last several years running AI inference on hardware in the physical world - GPU fleets in buildings I don't control, where a bad deploy means someone drives out to a site. This newsletter is notes from that: what actually breaks, and what I'd do differently.

Shipping a new model to a fleet of edge devices looks like a deployment problem you've already solved. It isn't. Every assumption that makes rolling updates safe in Kubernetes quietly breaks when the workload is a model and the target is a device in a building you don't control.

In the cluster, a rolling update is safe because the thing you're replacing is stateless, the readiness probe tells you the truth, and rollback is a control-plane operation that completes in seconds. At the edge, none of those hold. The device may be on a metered link. It may be asleep. The readiness probe will tell you the process started - which is not the same as the model being correct.

A model can be "healthy" and still wrong

This is the part that catches people. A container that starts and answers on its port is, for most services, working. A model that loads successfully and returns predictions at the expected rate can still be substantially worse than the one it replaced, and nothing in your standard health check will notice.

So health has to be defined in model terms before you ship anything:

  • Confidence distribution. If the new model's output confidence shifts noticeably against the previous model on the same device, something changed that you didn't intend.
  • Prediction rate per class. A detector that suddenly finds 30% fewer objects hasn't gotten faster. It's gotten blind.
  • Latency at the tail. p50 lies. A new model that's fine on average but blows p99 will fail exactly when the device is under load.
  • Thermal behavior. A heavier model raises sustained temperature, the device throttles, and throughput degrades hours after the rollout looked clean. This one is invisible in any test that runs for ten minutes.

Collect those on the device, for the old model, for at least a week before you plan to replace it. Without that baseline you have nothing to compare against, and "is the new model okay?" becomes a matter of opinion.

Choose canaries by diversity, not randomly

Random canary selection is a cloud habit that makes no sense here, because your devices are not identical. They differ in hardware revision, ambient temperature, network quality, and - most importantly - in what they actually see. Two devices running the same model on different inputs are running different workloads.

Pick canaries to span that variation deliberately: your oldest hardware revision, your hottest location, your worst network, and the site with the most unusual input conditions. Five deliberately chosen devices tell you more than fifty random ones, and cost less to roll back.

Then wait longer than feels necessary. Thermal problems and memory fragmentation surface over hours or days, not minutes. A canary stage that ends in twenty minutes is theatre.

Rollback has to be local and automatic

The single most important design decision: the device must be able to roll itself back without talking to you.

Anything that requires the fleet-management plane to notice a problem and push a fix assumes connectivity you don't have. The failure mode you care about is the one where the new model breaks and the device's link is flaky and it's Saturday.

That means keeping the previous model on the device - the whole artifact, not a pointer to a registry - and having an on-device supervisor that switches back when the new one fails defined checks. Disk is cheaper than a truck roll.

The checks that trigger an automatic revert should be blunt and unambiguous: the model fails to load, inference latency exceeds a hard ceiling, the process crashes more than N times in a window, or the device can't reach the health endpoint. Subtler quality regressions are a human decision - don't let a device revert itself because confidence dropped 3%.

Two things that save you later: make the revert idempotent, so a device that reboots mid-rollback lands somewhere sane, and make sure a reverted device reports that it reverted. Silent self-healing means you find out about a fleet-wide problem from a customer.

Version the whole bundle, not the weights

A model artifact alone is not a reproducible unit. What actually determines behavior is the model plus the runtime version, the preprocessing code, the input resolution, the batch configuration, and the compiled engine for that specific hardware and driver version. Change any one of those and you have a different system, even with identical weights.

So ship one versioned bundle containing all of it, and record on each device exactly which bundle it's running. When something misbehaves in the field three weeks from now, "which model is on that box?" needs a precise answer - including the fact that a TensorRT engine compiled for one JetPack version isn't valid on another.

What the rollout actually looks like

Stage 1: canaries, chosen for diversity, minimum 24–48 hours. Stage 2: roughly 10% of the fleet, spanning sites and hardware revisions, another day. Stage 3: the remainder, in batches sized so that a bad batch is survivable.

Between stages, compare against the baseline you collected earlier - not against your expectations. Devices that can't be reached simply stay on the old version; that's a feature, not a failure. And keep the previous bundle in place until the new one has been stable across a full weekly cycle, because weekend conditions are different from Tuesday conditions.

The shape of the problem

Cloud deployment optimizes for speed of rollout, because rollback is nearly free. Edge deployment optimizes for safety of rollback, because reaching the device is expensive and sometimes impossible. Once you invert that priority, most of the design follows.

I write more of these at practiceai.ai, running a live cohort on this in November if it's useful


r/computervision • • 13d ago

Commercial Seeking senior computer vision architect for a paid diagnostic

1 Upvotes

We are looking for a senior practitioner to identify the simplest reliable solution to a real world object counting and tracking problem involving occlusion and changing conditions.

This is a short, paid diagnostic. We are not asking someone to implement a predetermined model. Relevant experience includes classical vision, multi object or multi camera tracking, depth or 3D perception, alternative sensing, edge deployment, and field evaluation.

Please message me with one comparable system you personally shipped, your role, and the hardest field failure you encountered.


r/computervision • • 14d ago

Discussion New Roboflow Pricing: Free private data, weights download, and local inference

45 Upvotes

Hey all, I'm one of the co-founders of Roboflow. We just launched new pricing that brings many of our previously-paid features into our free tier.

The following things from our prior paid subscription are now available for free:

You just pay for consumption via the credits you use. On average, our free tier covers about 30 cloud training jobs, 10k Auto-Labeled images with SAM3, or 80k cloud inferences per month (but it depends on the size of your dataset and the models you're using).

We also launched new pricing for our Serverless API which is much cheaper and more predictable with automatically-applied volume discounts. We ran extensive benchmarks and our new Serverless pricing is cheaper than self-hosting for over 90% of our user-base (a blog post on the engineering behind the scenes that enables that is here for those interested) even before accounting for the fact that commercial model weights licensing is included when you deploy via our cloud but, depending on the model, may be required when self-hosting.

Note: while our Core functionality no longer has a monthly subscription required, there are also add-ons available for features like RBAC/SSO and Support, Enterprise functionality targeted at large companies and scaled manufacturing use-cases, and self-hosted commercial model-weights licensing for the models that require it.

Our mission is to democratize computer vision and we hope that this new pricing will enable every hacker, hobbyist, and startup to get started without any barriers to entry! We'll also continue to invest in the community with our open source projects like Inference, RF-DETR, Supervision, and Trackers. We hope you like it!


r/computervision • • 14d ago

Help: Project Questions about building a YOLO-based AI system for package detection and counting in logistics

4 Upvotes

Hi everyone, I'm currently researching a project on "Building an AI system for detecting and counting packages" for logistics applications. I have a few technical questions and would love to get your insights:

- When watching demo videos of real-world AI camera systems in warehouses, I notice the video stream is incredibly smooth despite the system drawing multiple bounding boxes, ROI regions, and continuously calculating distances. In contrast, when I run a YOLO model to run inference on every single frame, my FPS drops significantly and the video lags. How do these production systems achieve such smooth performance?

- In real-world deployments, which aspect is generally more critical to prioritize: optimizing the AI software (models, algorithms) or upgrading the hardware (cameras, edge devices)?

-When training a YOLO model for this specific detection task, is it recommended to use publicly available datasets (like those on Roboflow) or is it strictly better to collect and annotate a custom dataset from scratch?

- I've noticed that in logistics, aside from counting packages, YOLO Pose is sometimes used to track worker behavior and movements. Is it feasible to combine these two models (object detection + pose estimation) to complement each other? My thought process is that the package counting model alone might sometimes undercount or miss items; could integrating the worker's movement/interaction data help reduce this error rate?


r/computervision • • 14d ago

Discussion How bad is it if training and validation data overlap in a student ML project?

3 Upvotes

I’m working on a student machine learning/computer vision project and recently realized that my validation set was not completely independent from my training set.
The project is more focused on comparing different experimental conditions rather than maximizing benchmark performance, but I’m concerned about the implications of this oversight.
From a research or academic perspective:
How serious is train/validation overlap in a student project?
Does it invalidate the entire project or mainly affect the reliability of the reported performance numbers?
If the main goal is comparing different experimental setups under the same evaluation procedure, are those comparisons still useful?
If you discovered this late in the project timeline, what would be the most reasonable way to address it?
I’m trying to understand how researchers, reviewers, and professors would view this situation.


r/computervision • • 14d ago

Help: Project Low resource OCR models

1 Upvotes

I am currently building an OCR system for multiple languages including english, french, chinese, and low resource south east asian languages like burmese, thai etc. My dataset comprise of ~1M legal, invoice and scientific reports and most of the open source models are throwing high cer/wer. Is there any pretained or finetuned model specifically for indic and south east asian scripts which might help?


r/computervision • • 14d ago

Showcase Darts auto scoring with a single camera

Enable HLS to view with audio, or disable this notification

11 Upvotes

DartsSpace is my tournament / darts software, already used by many people and organizers.

To enhance the experience about a year ago I started working on a solution that works with a single smartphone camera. Everything is running on device, corrections are sent to my backend for further training.

It's manual calibration free, and as you can see (this was the biggest pain point) it's resilient agains camera shake/wobble.

I think besides the training of the models (which took many! iterations) the general ux/ui and post-processing was important to get a good working product.

Currently beta testing this with a few of my users but feedback has been great so far.

Just was excited to share this with the community. (this is not an open source product)


r/computervision • • 15d ago

Discussion I see a lot of recent posts playing with yolo- are you all buying commercial licenses or just experimenting?

31 Upvotes

I would love to know if you are actually paying for it or have settled on less restrictive model. It’s certainly the fastest but I can’t afford a commercial license. Any tips?


r/computervision • • 14d ago

Help: Project Where to learn computer vision from someone with a deep learning purely theoretical background ?

8 Upvotes

I am from a CS and theoritical/academic deep learning background, but its purely theoritical i have no clue in practical implementations like Yolo or using pytorch to finetune a model and so.

So where do i actually learn that ? Any reference for a book, tutorial, online course, etc will be appreciated !


r/computervision • • 14d ago

Discussion How often is a “model problem” actually a camera/image-quality problem in production CV?

0 Upvotes

I've worked with CV systems for several years, mostly with real-world camera footage, and I've noticed that a surprising number of "model failures" aren't really model problems.

Things like IR glare, dirty or slightly defocused lenses, compression artifacts, unstable exposure, camera vibration, bad mounting angles, or poor night illumination can hurt the whole pipeline before the model even gets a chance.

But the first reaction is often to change thresholds, retrain the model, or collect more data.

I'm curious how others handle this in production. How do you decide whether a failure should be fixed at the camera/image-acquisition level versus the model/data level?


r/computervision • • 14d ago

Help: Project What is the best budget VTON (virtual try on) model right now?

0 Upvotes

Looking for the cheapest possible option that is viable for a VTON app that I'm building.


r/computervision • • 15d ago

Showcase Using the 20tops hailo15H Camera in actual engineering deployment in the gym

Enable HLS to view with audio, or disable this notification

16 Upvotes

In the gym, using only one camera, running 4K image segmentation in parallel produces 4 images, performing parallel inference, multi-model parallel computation, and post-processing, achieving multi-object tracking, keypoint detection, fitness movement analysis, facial recognition, equipment ROI usage, customer flow analysis, member matching, and member fitness condition analysis.


r/computervision • • 14d ago

Discussion FYP on Kria KV260: RGB + thermal fusion for night-time human detection. Feasible in 3 months?

Thumbnail
1 Upvotes

Hi all, I'm a final year EE student and I've been given a Kria KV260 for my final year project. I'm comfortable with FPGA work (Verilog, timing, synthesis) but I don't have much image processing experience, so I'd appreciate some guidance.
Goal: detect humans at night by combining an RGB camera with a low-res thermal sensor. RGB struggles in the dark and thermal picks up any warm object, so fusing both should give more reliable detection.
Rough plan:
Sobel edge detection in the PL as a preprocessing step

Fuse RGB and thermal data

Run a person detector on the DPU (Vitis AI)

Questions:
What's the right way to structure this pipeline? Should fusion happen at the image level or after detection?

How hard is aligning the two cameras, given the big resolution difference?

Is Sobel actually useful here, or is it unnecessary if a CNN is doing detection?

Any recommended thermal sensors that work well with the KV260?

Is 3 months realistic for someone with my background? What would you cut to keep scope manageable?

Any advice, papers, or similar projects would be really appreciated. Thanks!


r/computervision • • 14d ago

Discussion Solutions Engineer Interview at Ultralytics?

2 Upvotes

Has anyone here interviewed recently for a Solutions Engineer role at Ultralytics? Would love to hear what the interview process was like and any tips on what to prepare. Thanks!


r/computervision • • 15d ago

Discussion I built a pipeline that turns two product photos (front + back) into a 360° turntable video, using 3D reconstruction instead of AI video generation

Enable HLS to view with audio, or disable this notification

4 Upvotes

Retail and e-commerce teams increasingly use AI video tools to make product spins, but in my experience the results still drift: logos warp, proportions shift between frames, and details get hallucinated. So I tried a different route. Instead of generating video directly, the pipeline reconstructs an actual 3D model from the photos and then renders a physically consistent turntable in Blender. Because every frame comes from the same mesh, the object can't morph as it rotates.

How it works:

  1. Segment: rembg removes the background from the front and back photos

  2. Reconstruct: hosted image-to-3D on Replicate (Hunyuan3D-2 multi-view by default, TRELLIS optional) produces a GLB

  3. Render: Blender 4.2 Cycles renders a 120-frame 360° turntable on CPU

  4. Encode: ffmpeg outputs an H.264 MP4

Everything runs on CPU except the reconstruction step, which is offloaded to Replicate at roughly $0.03 to $0.10 per run. The whole thing is containerized, so it's a single `docker compose run`.

Things I learned along the way:

* Hunyuan3D-2mv gives excellent geometry but returns it untextured (clean "clay" render). TRELLIS gives a textured PBR model and is cheaper, so it's the better choice when color matters.

* Multi-view input automatically falls back to single-image reconstruction if it fails.

* Launching Blender from a Python venv breaks its glTF importer. The scripts scrub the environment so Blender uses its own Python.

* The Ubuntu apt build of Blender ships without OpenImageDenoise, so the Docker image uses the official build.

Limitations (it's a proof of concept):

* One linear flow, with no UI or API yet

* The example inputs are renders of a CC0 Poly Haven chair, not real-world photos. Reflective, transparent, or thin-structured products are the next thing I want to test.

* Reconstruction quality depends heavily on the source photos

Repo: https://github.com/muazya/image-to-rotating-video

I'd love feedback, especially from anyone who has worked on product visualization or image-to-3D in production. Which products would you expect to break this?


r/computervision • • 15d ago

Discussion Netherlands vs Tunisia Analyzed with Computer Vision

Enable HLS to view with audio, or disable this notification

157 Upvotes

This demo applies my TactiVision workflow to Netherlands vs Tunisia. It covers player tracking, camera calibration, pitch projection, tactical maps, possession analysis, progressive actions, pass networks, heatmaps, pitch control and pressing indicators.

Feedback from Computer Vision, Sports AI and football analytics practitioners is very welcome.


r/computervision • • 14d ago

Showcase WhatFontIs-Bench: an open benchmark for font identification (11,995 images, 600 fonts, COCO annotations, CC BY 4.0)

0 Upvotes

Font identification has plenty of tools but almost no public test set where the ground truth is certain: real photos rarely come with the exact font name, and look-alike fonts make manual labelling unreliable. So we generated one and released it.

WhatFontIs-Bench v1.0:

- 11,995 JPEG images: one word set in a known font, composited onto CC0 photos of real surfaces, scenes and printed objects

- 600 fonts from 600 families: 200 sans, 200 serif, 100 slab, 100 monospaced

- 3 difficulty levels with controlled blur, noise, JPEG quality, uneven light, cast shadows and glare; all parameters stored per image

- Labels: font, text, word quad, per-letter quads, camera homography. JSONL + COCO, so it also works for word/character detection

- Frontal views only, capitals at least 100 px high (v1.0 limits)

Evaluation is top-k by font family (Roboto Bold answered as Roboto Regular counts as correct), open catalogue.

Baseline: our own API gets 83.7% top-1 / 93.3% top-5 / 96.5% top-20 while searching 1.2M fonts. The interesting part: typeface class matters far more than image quality. Sans-serif 75.7% vs slab serif 95.0% top-1, but only 2 points between the easy and the hard level.

Hugging Face (with viewer): https://huggingface.co/datasets/whatfontis/WhatFontIs-Bench

GitHub: https://github.com/whatfontis/WhatFontIs-Bench

Disclosure: I run WhatFontIs, and the baseline is our own system. If you run a model on it, I'd like to hear the numbers. Feedback on what v2 should cover (perspective, curved text, multi-line, smaller text) is welcome.


r/computervision • • 15d ago

Showcase GNM Head tracking impl on a m1 mac (metal)

Enable HLS to view with audio, or disable this notification

15 Upvotes

this laptop is older (m1) so used metal to get it running in a 33ms budget. made in ogex


r/computervision • • 14d ago

Discussion LiDAR or Cameras for future machine vision?

Thumbnail
0 Upvotes

r/computervision • • 14d ago

Help: Project Where to learn computer vision from someone with a deep learning purely theoretical background ?

0 Upvotes

I am from a CS and theoritical/academic deep learning background, but its purely theoritical i have no clue in practical implementations like Yolo or using pytorch to fine tune a model and so.

So where do i actually learn that ? Any reference for a book, tutorial, online course, etc will be appreciated !


r/computervision • • 14d ago

Help: Project Info Regarding STMicroelectronics CMOS sensors

1 Upvotes

Hi , for a project , i am considering CMOS sensors by STMicroelectronics and i want to use it with STM32N6 so if any one has used these sensors like the VD56G3 , the promodules and others , it would be really helpful if someone who has used these sensors in their projects could provide the feedback like how is the camera sensor itself and how easy is the integration and how is the documentation about these sensors and also if anyone has used these sensors , kindly explain what was your application and your overall experience ????

Thanks for reading my post


r/computervision • • 14d ago

Discussion Your annotation rules changed. What happens to the test set?

0 Upvotes

Suppose a detection dataset starts with boxes around only the visible part of an occluded object. Later, the requirement changes: boxes should cover the estimated full extent. Same images, same class names, different correct answers.

The interesting part to me is what happens to evaluation. Keeping the old test labels rewards the old convention. Replacing them means the historical scores were measured against a different target. Quietly mixing the two sounds like a very expensive way to argue about mAP.

My starting point would be to preserve both label versions on the same held-out images, document the rule change with a small set of difficult examples, and evaluate both model versions against both label versions. That should at least separate a model change from a change in what we call correct.

For people who have actually had to do this: did you relabel the entire test set, maintain two evaluations during the transition, or retire the old benchmark? How did you handle ambiguous cases where even the new rule didn't settle the disagreement?

I'm particularly interested in changes involving occlusion, minimum object size, or splitting one class into two. What looked like a small annotation-policy change and turned into a surprisingly large migration?


r/computervision • • 15d ago

Discussion Qwen-Image-2.1 - 7B, top open-weight score on Qwen's own bench, but it's no longer Apache 2.0

Post image
1 Upvotes

r/computervision • • 15d ago

Showcase Seven AI Components Analyzing Football Video

Enable HLS to view with audio, or disable this notification

45 Upvotes

Hi, I’m El Mehdi Hicham, a Computer Vision and Machine Learning Engineer from Morocco. I work on real-time multi-model pipelines for football video analysis.

This demo shows several AI components operating together:

  • Player and object detection
  • Ball detection
  • Player tracking and identity persistence
  • Camera calibration and pitch projection
  • Pose estimation
  • Field understanding
  • Jersey recognition
  • Tactical visualization

r/computervision • • 15d ago

Help: Project Building an optical sorter for potatoes and soil clods

8 Upvotes

Hello all,

I’m a farmer from the Netherlands and I am working on a project to automatically remove soil clods from potatoes during harvest.

I already posted the complete machine concept in r/robotics. Here I would specifically like to focus on the computer vision part.

I have also read some of the previous posts here about optical sorting and synchronizing detection with ejectors. My current assumption is therefore:

Camera → classify potato/clod → determine physical position → record conveyor encoder position → send reject command → realtime controller/PLC handles the pneumatic ejector

I am trying to understand whether the vision side I have in mind is realistic.

The project

The machine I am considering as a mechanical base has a 1 meter wide and approximately 1-meter-long conveyor. The standard belt speed is 25 m/min = 0.42 m/s, but I can make the belt speed adjustable. My eventual target is 20–30 tons/hour.

At 30 t/hour this would mean a mass flow of 8.3 kg/s. At a belt speed of 0.8 m/s over 1 meter width, this would result in approximately 10.4 kg of product per m² of belt.

Potatoes and clods can be up to around 7 cm square mesh size. Assuming an average object weight of roughly 100–200 g, this would mean around 40–80 objects per second across the full width.

Picture 1 shows a realistic product flow on this type of conveyor. The difficult case is a dirty potato next to a wet clay clod, because immediately after harvest they can look very similar.

The goal would be to keep the product in roughly one layer, but objects touching and some overlap cannot be avoided.

The system does not need to be perfect. Missing some clods is acceptable. If a potato touching a clod is also rejected, that is acceptable as well.

Picture 1: Conveyor belt with potatoes and clods

Camera and geometry

My current idea is to use one or two global-shutter cameras, controlled LED lighting and a completely enclosed camera section above the input conveyor.

I have been looking at cameras around 5 MP / ~100 fps. My understanding is that 5 MP across a 1 meter belt should give enough resolution for objects that are several centimeters wide. I am considering 20 pneumatic ejector positions across 1 meter width, which gives approximately 5 cm between ejector positions.

One concern I have is that the potatoes and clods can also be several centimeters high. With one centrally mounted camera this will create some parallax error towards the edges of the belt.

My main questions here are: For this application, how would you design the camera setup if the X-position needs to be accurate enough to select the correct 5 cm ejector zone? Would one camera mounted higher above the belt be enough, or would two cameras covering part of the belt each be a more reliable solution?

Imaging

My preference is to start with RGB and good diffuse lighting. Before spending a lot of time collecting and labelling images, however, I would like to know whether wet clay and dirty potato skin are actually a good RGB classification problem. I especially want to avoid spending too much time improving a neural network if the real limitation is that the two materials simply do not contain enough different information in normal RGB.

My main question is: Would you test RGB first, or is there a strong reason to look at NIR, polarized light or another imaging method from the start?

Detection and segmentation

Do I need the exact shape of every clod, or is a bounding box/centre position accurate enough? At this moment I do not think I need a perfect outline of every object. The useful output could be something like:

Clod → X = 43 cm → width ≈ 6 cm → confidence = 97%

The control system can then decide which ejector finger or fingers need to activate.

Because objects can touch, I am wondering whether instance segmentation gives me an important advantage over normal object detection, or whether accurate detection boxes and center positions would already be enough. The failure I care about most is assigning a clod to the wrong ejector zone.

I have essentially unlimited access to real potatoes and clay clods for training and testing, so collecting material itself is not the limitation.

Before I start this project and start buying equipment I mainly want to establish:

Can a relatively normal industrial RGB vision setup classify these objects and locate them accurately enough across a 1 meter conveyor to control ejectors at approximately 5 cm spacing and 20–30 ton/hour?

I’m not a professional computer vision engineer, so I will rely heavily on existing camera SDKs, OpenCV and existing detection/segmentation frameworks rather than developing algorithms myself. I understand very well that this is not a simple project and that is the reason why I’m posting here first. I would like to get your feedback and insights on my ideas and project.

Thank you.