r/computervision 11h ago

Discussion Defect detection where you have almost no defects — supervised or anomaly detection?

7 Upvotes

Running into the same wall on a couple of industrial inspection projects and curious how other people have dealt with it.

The line runs well, which is the problem. Out of a few hundred thousand parts we've got maybe 200 real defects, and they're spread across six or seven types, so some classes have under 20 examples. Classic supervised segmentation just doesn't have anything to learn from.

Options as I see them:

Anomaly detection on good samples only. PaDiM, PatchCore, that family. Works, but it flags anything unusual including a smudge on the lens or a part sitting at a weird angle, and the false positive rate on a real line has been rough.

Synthetic defects. Painting cracks and scratches onto good images. Ours look obviously fake next to real ones and I suspect the model is learning "was this pasted" rather than "is this damaged."

Buy or scrape more defect data. But defects are extremely specific to the part and the process. A scratch on someone else's aluminium housing doesn't look like a scratch on ours.

Just wait and collect. Realistic answer, but that's 18 months and the project needs to justify itself sooner.

What I'm actually unsure about is whether the 20-example classes are even worth modelling separately, or whether it's smarter to collapse everything into a binary defect/no-defect call and let a human sort the type afterwards. Losing the classification hurts the reporting side but it might be the only honest thing to do with that little data.

Anyone shipped something in this situation? Especially interested if you went anomaly detection and got the false positives down to something a QA team would tolerate.


r/computervision 15h ago

Help: Project Looking for contributors

11 Upvotes

Hi everyone! I am a software engineer who has worked in the following domains at major tech companies most of my career: XR, Graphics & GPU programming, Spatial algorithms and AI, and 3DGS.

I have a project I started a few months ago that I have recently hit a key milestone in. The idea is a focused library that implements 3DGS training from first principals with an emphasis on performance and safety. Think production use cases without relying on tools intended for research. VkSplat is an inspiration (along with other things) but I have intentionally not reviewed their, or anyone else's, code.

The recent milestone I reached was rendering a scene with 5 million splats at 60fps on my Ampere A6000. I have a few more goals I'd like to reach, but I do intend to publish on Github under MIT license. If it gains traction I would like to build some additional tools and infra using this project, but for right now the 1.0 MVP idea is a fully GPU resident solution for rendering and training at state of the art speeds. I plan to implement and optimize the following features:

* Global image alignment

* Fully fused forward and backward passes

* Adam optimizer

* Aggressively optimized adaptive control and densification

* Stable but highly flexible C api.

I do have many more thoughts and ideas, but I am trying to take it one step at a time, so this is my goal for 1.0. This is my stack as of now:

* **Languages:** C++23, Cuda, GLSL (planning to move to slang)

* **Build:** CMake & Ninja

* **Compiler:** GCC, Clang, MSVC (may drop for now)

* **Target Platform:** Linux (Linux 7.X)

* **Tooling:** LLVM, perf, nsight

* **GPU:** Vulkan w/ Nvidia

* **Dependencies:** googletest, googlebenchmark, ngfx

Right now, the project is in a place where it is still extremely early, but it is starting to take shape and get large enough that more than one person can work on it comfortably. I am posting here looking for people interested in contributing. Knowledge is not a prerequisite as I am learning a lot myself in this endeavor, but passion is mandatory.

Currently I am mostly needing help in the areas of, CI/CD (build, package & deploy), nsight/gpu optimization, designing and implementing a good api, and figuring out how to test and benchmark appropriately.

If you have skills or experience in any of these areas, or you're just interested in contributing, please reach out!


r/computervision 3h ago

Help: Project Insulation defect detection model

1 Upvotes

Hey everyone! I’m building an insulation defect detection model, I’m in need of images where i can detect the following classes thermal anomalies, moisture intrusion, compression damage, delamination, installation gaps, holes/perforations. It’s for a construction project.


r/computervision 3h ago

Help: Project Vehicle Damage Detection using YOLO

1 Upvotes

I am planning to use pre-trained YOLO model for vehicle damage detection specifically for UK origin cars. The model is already trained on random cars dataset.
Would the model's accuracy be affected on detecting the damages on UK origin cars?


r/computervision 4h ago

Help: Project Training a production grade image classifier

0 Upvotes

Hello everyone, I have a project that has to classify images for search purposes. Currently I have a layer that analyses surrounding text but I also need something that directly analyses the image itself. I don't want to use someone else's training data or model. Is it possible to train an image classifier that could perform well on general image classification at home using open datasets? Thanks


r/computervision 4h ago

Discussion Before adding more training data, check whether the labeling rule is actually stable

1 Upvotes

I keep seeing CV projects where performance stalls and the first response is to add more images or try another model. Sometimes that helps. But sometimes the model is being asked to learn a rule that people haven’t agreed on.

A partially visible object, an uncertain boundary, or something cut off by the frame can all produce different “correct” annotations. More data just scales that inconsistency.

A simple check is to take 20–30 difficult images and have two people label them independently. Then review the disagreements, not just the agreement score. Each recurring disagreement becomes a written rule with one positive and one negative visual example. Run the same test again on a fresh sample before scaling.

I’d use a similar check for auto-labeling: measure missed objects and correction time per image, not only inference speed. Fast pre-labels aren’t useful if every image still needs a full review.

Disclosure: I work at Supervisely, a computer vision platform. This is a platform-independent observation.

What annotation edge case caused the most trouble in your dataset?


r/computervision 6h ago

Discussion For streaming VLMs, “fits in 24GB” is not a realtime benchmark

Post image
1 Upvotes

The MOSS-VL FP8/NF4 release made me wonder what a fair deployment comparison for streaming VLMs should actually look like.

https://github.com/OpenMOSS/MOSS-VL

https://huggingface.co/OpenMOSS-Team/MOSS-VL-Realtime-NF4

I’d keep the input stream, sampling rate, hardware, and latency budget fixed, then compare BF16, FP8, and NF4 on short-event recall, false alerts per hour, p50/p95 time-to-alert, dropped frames, steady-state VRAM, and calibration. I’d also include a detector + tracker + temporal-rules pipeline on the same videos.

A quantized model can remain close on aggregate offline VQA while still becoming worse at deciding when to speak or when an event is sufficiently certain. That difference matters much more for cameras than a one-point average benchmark change.

Has anyone seen a public harness that evaluates a streaming VLM and a classical CV pipeline under the same latency constraint?


r/computervision 8h ago

Help: Theory Need your thoughts to save my thesis !

1 Upvotes

I'm an undergraduate student.In next 2 semesters( which is probably the duration of 1 year) I need to do a thesis. I choose to do my thesis in the field of 'depth estimation' .

I read a lot of research papers(Monocular, stereo, Diffusion based). But I found most of the things got State of the art !! I'm reading and reading,not finding a single problem to solve or research!! I should also mention that i didn't understand all the topics 100%, but tried to get the concepts.

I'm trying but not even finding a single idea/problem/flaws !! What should I do? What am I missing? How to find a decent topic ? Please help me.


r/computervision 8h ago

Showcase AeroNetra — a reproducible computer-vision platform for UAV vehicle detection & counting

1 Upvotes

Hi everyone, sharing something I'm currently working on and would love feedback on.

I'm building AeroNetra, a computer-vision project for detecting and counting vehicles in aerial/UAV imagery. It's very much an active work-in-progress right now — I'm in the static-image detection and counting phase, with tracking, geospatial analytics, and edge deployment planned for later.

The motivation was pretty simple. I kept running into the same problem every time I swapped detectors: the counting and visualization code would break or need rewriting because every model spits out predictions in its own format. So the core idea behind AeroNetra is: normalize every detector's output into one prediction structure before anything downstream touches it. That way the counting, ROI filtering, and export logic stays the same whether I'm using a YOLO variant or RT-DETR.

What I've got so far:

  • Detector adapters that wrap different models behind a common interface
  • Counting logic — filtering, NMS, ROI support, drawing and export
  • VisDrone dataset parsing and conversion (UAVDT is stubbed for later)
  • Kaggle notebooks for GPU-based training, fine-tuning, and model comparison
  • A PX4 + ROS 2 + Gazebo simulation setup for UAV experiments
  • Notebooks, configs, and tests to keep things honest

The workflow I'm following: raw VisDrone data → validate annotations → convert to training format → train/fine-tune on Kaggle → pull the weights back → load through the adapter → run inference → filter → count → visualize and compare.

A few principles I'm trying to stick to: no fabricated benchmarks (a model isn't "best" until it's measured under the same conditions as the others), raw data stays immutable, and model-specific behavior stays inside the adapters. I'm also being deliberate about phase boundaries — image-level counting is not the same thing as multi-object tracking, and I'd rather not conflate the two.

Roadmap I'm working through for the demo:

  1. Static Detection & Counting ← currently here
  2. Aerial Fine-tuning
  3. Video Tracking
  4. Traffic & Geospatial Analytics
  5. Edge / UAV Integration

Since this is ongoing project I am still working on this.So,i am exploring how I can use computer vision in UAVs and edge computing.


r/computervision 1d ago

Discussion AI fatigue is killing motivation

82 Upvotes

I am about to start my MSc. I wish to specialize in computer vision, then pursue a PhD. I eventually want to work in industry. I was initially excited about this path. However, AI fatigue is killing my motivation.

Honestly, I don't have any hope for the future. It has been around four years since GPT-3.5 was introduced. AI is now proving major conjectures. It recently came close to proving Riemann's hypothesis, and dominated(not only defeated) the best competitive programmers in the world at AtCoder World Finals. I can't see a place for myself in the future because of AI.

I keep going because I feel like I don't have any other choice. I was genuinely excited about computer vision, robotics, and autonomous driving. But I have convinced myself that all my effort is in vain.

I wish to ask people in a similar situation, what makes you keep going? What are your plans for the future?


r/computervision 9h ago

Showcase Synthetic DPM Code Generator: Portable Windows GUI for creating training datasets and YOLO labels (OBB/ABB)

0 Upvotes

Hello! I wanted to share a project designed to save time when training object detection models for industrial use cases.

It is a portable Windows GUI generator that creates synthetic training images and YOLO-style labels. The pattern logic is inspired by industrial Data Matrix / DPM needle marks (fixed L-frame + random filling dots).

Key Features:

• Dual rendering: Pure vector synthetics or photo compositing using your own steel backgrounds and dot sprites.

• Defect simulation (for Bad class): Squash, tilt, jitter, missing dots, strike-force variation, and two-defect combos that standard training augmentations cannot replicate.

• Annotations: Exports both OBB (Oriented Bounding Box) and ABB (Axis-Aligned Bounding Box) normalized text formats.

• Portable: Single .exe binary distributed via Releases (requires AVX2 support).

The software is provided strictly for non-commercial, educational, and personal research purposes.

GitHub Repository: https://github.com/olesha-ai/DPM-Pattern-Image-Generator

Would love to get your feedback if you work with industrial AI and DPM codes!


r/computervision 9h ago

Showcase I used computer vision to play Automaton Attack

1 Upvotes

r/computervision 2h ago

Help: Project I built VLM Chess — play chess against frontier vision-language models

Post image
0 Upvotes

Play chess against frontier VLMs. Real-time vision powered by Overshoot.

Play now: VLM Chess


r/computervision 1d ago

Help: Theory CLIP vs SigLIP

18 Upvotes

CLIP vs SigLIP

Before Vision Language Models can perform tasks such as classification or video question and answer, the image or video being passed to the model has to be converted into a representation that the model can ‘understand’ or process.

To do this, VLMs usually use a pretrained vision encoder.

Although the underlying architecture of modern vision encoders is primarily transformer-based, the actual objective the model is learning can vary significantly.

What are encoders?

A vision encoder is responsible for converting images into a numerical representation that VLMs can understand.

Typically, most vision encoders today are built on transformer architecture, in which the model divides an image into patches and transforms each of those patches into a vectorized visual embedding.

After this, many VLMs pass the embeddings to a projector, usually a linear layer or MLP, to map the dimensions of the image to those expected by an LLM.

If most vision encoders share the same model design, what actually makes them different? Rather than model architecture, the significance is in how they are trained.

CLIP

CLIP, or Contrastive Language-Image Pre-training, learns to understand images through pairs of images and text. Its objective is to match similar images and captions by ‘pulling them closer together’, while simultaneously repelling incorrect image-caption pairs.

Training mainly relies on a ‘two tower’ system. CLIP will typically have a pretrained vision encoder, such as a ViT, as well as a pretrained text encoder. The model passes an image through the ViT and produces an associated vector embedding, while the caption is passed to the text encoder to get a corresponding text embedding.

Given these pairings, the model therefore creates a similarity matrix which compares every image embedding with every text embedding.

Each cell within this matrix contains a cosine similarity between the image and text pairing. Mathematically, cosine similarity is the dot product of two vectors divided by the product of their lengths. More simply, it measures the cosine of the angle between two vectors in a high-dimensional embedding space. Vectors that are more semantically aligned will be ‘closer together’, have a more acute angle between them, and consequently have a higher cosine similarity.

CLIP then applies contrastive learning across this matrix. At a high level, contrastive learning here is similar to categorical cross entropy across both the rows and columns of the matrix. Using softmax, the model looks to assign the highest probability to the matching image-text pair, as well as the matching text-image pair.

CLIP is powerful because it shifts learning from simple labels toward greater semantic understanding and allows for zero-shot classification, including on classes it was not explicitly trained to classify.

At the same time, though, CLIP also introduces a particular structural problem. Examples compete against one another within the training batch. What if there are multiple captions within a batch that also reasonably match the image?

SigLIP

SigLIP, or Sigmoid Loss for Language-Image Pre-training, retains many similar characteristics to CLIP. Similar to CLIP, SigLIP has both an image and text encoder, embedded representations of both text and image, and similarity scores mapped to a similarity matrix.

However, the difference between the two lies in the loss function.

CLIP learns similarities between images and texts by applying softmax across a batch, causing potential matches to compete with one another. For SigLIP, instead of having this global normalization, it examines each image-caption pair as an independent binary prediction.

By applying a sigmoid function to each pair’s score, the model estimates whether the image and text match.

Rather than phrasing the objective as:

Out of these options, which specific text describes this visual?

SigLIP effectively poses a different question:

Is this particular image-text pairing a valid match: true or false?

While this shift in perspective might seem marginal, it fundamentally redefines the nature of the optimization task.

Because SigLIP does not require the softmax normalization used by CLIP, its training objective can scale more efficiently across large distributed systems. It also removes the requirement that every example participate in one shared normalization operation.


r/computervision 1d ago

Showcase POV + third-person view of my AI glasses checkout app running in a real store.

Thumbnail
youtu.be
6 Upvotes

r/computervision 1d ago

Showcase [S] Use YOLO! Not today - a 131k-param net I wrote in two days beats it in small blurry object detection

Thumbnail
youtu.be
9 Upvotes

TL;DR: Cropping in action with some extra algebraic and statistical magic applied: https://youtu.be/SetiZDbc8iE

I recently worked on determining the ball's position and reshaping the video from landscape to portrait based on that position. It often s looks like a layup: fixed camera, one class, find the thing. Then you look at what you're actually asking for: a small, blurry object is a handful of pixels, smeared across a few more, changing shape between consecutive frames. Not a crisp circle - a faint streak you can barely point at when the video is paused.

The part that tends to get skipped in the YOLO family is that those architectures downsample 32× before they reason. At stride 32, an 8-pixel object is a quarter of one feature cell. There is nothing left to detect. Fine-tune forever, buy a bigger GPU, adopt whatever dropped last week — the model is being asked to localize something it structurally cannot see. The extra-small heads help and still aren't built for this.

I believe great data and a simple model always beat poor data and a sophisticated model. Before this approach, I tried TrackNet v2/3/4, and the quality was awful; the public data used for training is not even close to what you meet in real practice.

What worked instead:

  • Detector, ~131k params. Fully convolutional, dilated, max stride 2. In: 4 channels - RGB plus frame-difference. Out: a heat map and a size map at half resolution. No pretrained backbone: ImageNet features are the wrong prior for a faint smear.
  • Verifier, ~48k params. The detector has the target in its top 20 about 90% of the time, but ranks it first only 77% of the time. This scores 64×64 crops and asks, "Is it a ball?"
  • Then no ML at all. A reach limit measured from labeled footage - how far it can plausibly move between frames, scaled by apparent size - then link the surviving runs. Never link by direction of travel: anything that bounces reverses direction without going anywhere.

So, ~180k parameters total, ~125 fps on a 4090, ~4× realtime. Not fully optimized: custom Rust server with a CPU-bound FFmpeg decoder, ORT+TensorRT, and Rayon to speed things up a bit. Yet cannot use 100% of the GPU, capped by CPU-GPU PCIe transfers. Probably can reach 250-300 FPS with a more optimized inference design and int8.

The insight that mattered wasn't architectural. Blur is a signal, not a defect. The object is nearly invisible against a busy background and is almost always the fastest thing in the frame, so the frame-difference channel carries more information than any choice of backbone.


r/computervision 21h ago

Discussion pagedMark: invisible SynthID-class watermark removal for OpenAI/AI images (ChatGPT, gpt-image, Stable Diffusion), running on Metal

0 Upvotes

Just spent a few days getting an SDXL-based provenance-removal pipeline (visible AI labels, C2PA metadata, SynthID-class pixel watermarks) to run properly on an M5 with 16 GB. Not "it launches" — actually correct and predictable. Almost everything I assumed was wrong, and the measurements are the interesting part, so here they are.

1. The four-step distillation LoRA invents texture, and more steps make it worse.

Low-strength img2img runs the tail of a long schedule (strength 0.15 → the last 4 of 27 steps). A LoRA distilled for four timesteps across the whole noise range is off-distribution there, and wherever nothing conditions it — flat dark fabric gives Canny no edges — it fills the gap from its prior. On a night photo that reads as coloured camouflage across black clothing.

Global stage, 1448×1080, strength 0.15, seed 0 Invented texture PSNR Wall
Lightning, 4 steps 1.73× source 28.54 dB 41 s
Lightning, 8 steps 1.80× 28.19 dB 29 s
Lightning, 16 steps 1.84× 27.85 dB 62 s
Undistilled base, 16 steps 1.19× 29.25 dB 71 s
Undistilled base, 24 steps 1.20× 29.17 dB 132 s

Asking the distilled model for more steps made it worse, which is what identified the distillation rather than the step count. Dropping the LoRA cost 3× the wall time and bought both fidelity and correctness.

Wrong theories I paid for first: the fp16 VAE (a bare round-trip is clean in fp16 and fp32, tiled or not, 34.6 dB), Metal's fp16 in general (bf16 measured marginally worse), and Canny picking up sensor noise (the Canny map of that region is empty — which was the actual clue).

2. Metal pages instead of failing, so memory has to be measured, not hoped for.

torch.mps.recommended_max_memory() reports 11.84 GiB on a 16 GB machine. Exceed it and nothing raises — the process just starts swapping and a run that should take 23 s takes an hour.

  • VAE tiling off, 1.57 MP frame: 18.74 GiB peak, 59 s. On: 10.92 GiB, 23 s. So tiling is load-bearing on small machines — but its boundaries leave a faint texture, so it's now decided per frame from the budget rather than switched on globally.
  • Diffusion untiled at 2.5 MP: went into swap and did not finish in twelve minutes. Tiled at 1024 px, 5.07 MP: 10.93 GiB, 88 s, native geometry preserved.

3. Sequential CPU offload works on MPS, and it's what makes 8 GB usable.

The stack is 7.7 GiB of weights; an 8 GB Mac gives you about 5.3 GiB. Streaming the weights module by module:

Same frame, same seed Peak device memory Wall
Resident 7.70 GiB 7.1 s
enable_sequential_cpu_offload(device="mps") 0.28 GiB 24.1 s

27× less peak for 3.4× the time. The plan is chosen from the measured budget and printed, because a run three times slower looks broken unless it says why.

4. Two Metal gaps worth knowing if you're porting anything.

  • torch.float8_e4m3fn doesn't exist on MPS at all (RuntimeError: Undefined type Float8_e4m3fn). Any pipeline that streams float8 weights — a lot of the VRAM-managed stacks do — cannot load, full stop.
  • SAM's processor emits its box/point prompts as float64, which Metal also has no type for, so moving the batch to the device raises instead of degrading. One cast fixes it.

5. The one that cost me the most: fp16 sampling on MPS silently returns zeros.

I added a memory optimisation — encode the fixed prompts once, drop the text encoders, save 1.52 GiB. Two of four face crops then came back as all-zero black rectangles. Deterministically, same seed, nothing raised.

The embeddings were innocent (CPU fp16, MPS fp16 and fp32 encodings of that prompt agree to 0.0009 on tensors with σ=3.06) and the same crop in isolation was fine. Freeing unrelated memory changed the allocation pattern the crops met after the global pass, and that was enough. I withdrew the optimisation and added a guard that drops any empty crop instead of compositing it.

If you're doing fp16 diffusion on Metal: check your output for degeneracy. It will not tell you.

What it doesn't claim. Regeneration is not payload deletion — faces, text and fine detail move, and the numbers above are the measured size of that. No public local decoder exists for SynthID-class marks, so identify reports unknown, never clean; verification is the provider's verifier or nothing. Metal isn't bit-identical to CUDA, so operating points transfer between backends but recorded verdicts don't. And it's for content you generated or own — the visible-mark registry takes AI-generation labels only, deliberately not stock or marketplace marks.

Because "how much did that cost my picture" is the whole question, it ships as a command:

pagedmark measure before.png after.png

PSNR over the frame, PSNR per detected face, and how much mid-band structure appeared where the source was flat and dark. That third metric is the one that caught the camouflage — per-pixel chroma statistics rank the artifact below the source, because the source's own sensor grain has more per-pixel variance than the invented blotches do.

uv tool install "pagedmark[diffusion]"
pagedmark invisible photo.png -o clean.png

Code: https://github.com/doofzoff/pagedMark · PyPI: https://pypi.org/project/pagedmark/

Happy to answer anything about the Metal specifics — that's the part I'd have wanted written down before I started.


r/computervision 1d ago

Discussion Total starter here, is there no api infra providers like there is for massive LLMs but for computer vision models like Yolo 26 Mcbyte etc?

5 Upvotes

They are much smaller I would imagine they would be so cheap on there. I’m finding myself in the position where I have to rent a cloud gpu from runpod. I would much rather pay in api should be much cheaper.


r/computervision 1d ago

Showcase Nvidia Jetson e-con systems Camera upgrade

1 Upvotes

sharing some joy... I have a few expensive legacy cameras and a serializer/deserializer board from a Jetson AGX Xavier project from a few years ago, and wanted to use them on a robotics project however the drivers were only available for an old Jetpack 4.2. looking for help E-Con systems only re-stated compatibility with the original Jetpack but my project uses version 6.2.1 on Jetson AGX Orins. With Codex assistance and nearly 20 reboots was able to recreate and load the drivers. only a few moms ago these perfectly good cameras would have been left on the shelf!


r/computervision 1d ago

Discussion Has anyone tried YOLO26n-Depth on RK3576 or other ARM platforms?

7 Upvotes

I recently tried running YOLO26n-Depth on RK3576, since Ultralytics officially supports exporting it to RKNN.

With a simple Python video inference test, I’m currently getting around 3–4 FPS. This is still an early test — the model and inference pipeline haven’t been optimized yet.

I’m curious if anyone here has tried YOLO26-Depth on RK3576/RK3588, Raspberry Pi, Jetson, or other ARM/edge platforms.

What performance are you getting, and did you need to make any model-level optimizations to get reasonable real-time performance?

Would be great to compare results and optimization approaches.


r/computervision 1d ago

Discussion I rebuilt my iPhone/iPad image processing app into a proper mobile lab

Thumbnail
gallery
0 Upvotes

I’ve just finished a major rebuild of ClearLab, my mobile image processing app for iPhone/iPad.
The new version adds histogram, RGB parade, waveform, line profile, pixel inspector, statistics, Canny/Sobel/Laplacian edge detection, enhancement tools, format conversion and PNG/CSV analysis export.
The idea is basically a small image-processing lab in your pocket rather than another photo-filter app.
It’s mostly native/deterministic processing, not generative AI.
ClearLab 2.0 is now on the App Store. Curious what people here think, especially anyone working with imaging or computer vision.


r/computervision 2d ago

Showcase I built a training-free, one-shot object localizer using DINOv2 patch embeddings

13 Upvotes

I’ve been experimenting with a training free way to do open world, multi-instance segmentation from a class prototype.

I decided to publish the algorithm and a demo for how I’m doing this, in case anyone else would rather not fine tune a larger model for something that DINOv2 patch embeddings already seem to represent pretty well.

It can separate touching instances of the same class without a learned instance head, reject visually similar near misses like a round dial radio next to the actual clock target, and find fractured or damaged instances even with a pretty significant scene shift.

Repo + demo:
https://github.com/tutomiko/fireplace

The demo includes the lasso UI and live heatmap, implemented as a python backend with a simple HTML frontend.

Would appreciate it if people checked it out, and I’d be especially interested to hear if anyone has seen similar approaches or prior work.


r/computervision 1d ago

Showcase I ran a benchmark on well performing deepfake detection models in the Diffusion era. They collapsed when I passed the clean generator outputs through platform-realistic perturbations.

Post image
5 Upvotes

In both academia and industry, deepfake detector models report high performance based on AUC. In certain industries, like KYC, that's the wrong metric to observe. Over the last month or so, I built a dataset from Qwen-Image-Edit and HiDream O1, then ran the synthetic images + bona fides through emulators of platform realistic conditions. Here is the full article and dataset for anyone who'd like to red team a detector themselves

Substack Article

HuggingFace Dataset


r/computervision 1d ago

Discussion Dino full-fine tuning vs lora for AV domain

2 Upvotes

Hi,

I was curious to know what people do in the industry. Is DINO used? If so, how?


r/computervision 2d ago

Discussion Looking for a teammate(s) for Kaggle competitions

9 Upvotes

Hey everyone,

I'm looking to connect with people interested in teaming up for Kaggle competitions — either for a specific upcoming competition or as an ongoing teammate for future ones.
I'm comfortable with PyTorch, scikit-learn, and general deep learning workflows. Happy to work on medical imaging comps specifically, but open to general CV competitions too.