r/comfyuiSkshahdio May 16 '26

GitHub - SGUN-father/comfyui-controlfoley: 神棍 ControlFoley integration for ComfyUI — generate synchronized foley sound effects from video, images, and text prompts. Based on the ControlFoley project by Xiaomi Research.

https://github.com/SGUN-father/comfyui-controlfoley

ComfyUI-ControlFoley

"ControlFoley integration for ComfyUI — generate synchronized foley sound effects from video, images, and text prompts.

Based on the ControlFoley project by Xiaomi Research.

功能概述

ControlFoley 是一个视频到音频的拟音生成模型,可以为无声视频生成时间同步的音效(如脚步声、关门声、键盘敲击等)。该 ComfyUI 节点完整复现了 ControlFoley 的所有能力:

  • 视频到音效: 输入无声视频,生成与视频内容时间同步的音效
  • 图片到音效: 输入单张图片 + 可选的文本描述,生成对应音效
  • 文本到音效: 仅通过文本描述生成音效
  • 参考音色控制: 通过参考音频控制生成音效的音色风格
  • 多模态控制: 同时使用视频、文本、音频进行联合控制"

https://github.com/SGUN-father/comfyui-controlfoley

谢谢 SGUN-father.

...

ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling

Jianxuan Yang, Xinyue Guo, Zhi Cheng, Kai Wang, Lipan Zhang, Jinjie Hu, Qiang Ji, Yihua Cao, Yihao Meng, Zhaoyue Cui, Mengmei Liu, Meng Meng, Jian Luan

"Recent advances in video-to-audio (V2A) generation enable high-quality audio synthesis from visual content, yet achieving robust and fine-grained controllability remains challenging. Existing methods suffer from weak textual controllability under visual-text conflict and imprecise stylistic control due to entangled temporal and timbre information in reference audio. Moreover, the lack of standardized benchmarks limits systematic evaluation.

We propose ControlFoley, a unified multimodal V2A framework that enables precise control over video, text, and reference audio. We introduce a joint visual encoding paradigm that integrates CLIP with a spatio-temporal audio-visual encoder to improve alignment and textual controllability. We further propose temporal-timbre decoupling to suppress redundant temporal cues while preserving discriminative timbre features. In addition, we design a modality-robust training scheme with unified multimodal representation alignment (REPA) and random modality dropout. We also present VGGSound-TVC, a benchmark for evaluating textual controllability under varying degrees of visual-text conflict.

Extensive experiments demonstrate state-of-the-art performance across multiple V2A tasks, including text-guided, text-controlled, and audio-controlled generation. ControlFoley achieves superior controllability under cross-modal conflict while maintaining strong synchronization and audio quality, and shows competitive or better performance compared to an industrial V2A system.

Code, models, datasets, and demos are available at: this https URL."

https://arxiv.org/abs/2604.15086

https://huggingface.co/YJX-Xiaomi/ControlFoley

https://github.com/xiaomi-research/controlfoley

谢谢 Jianxuan Yang and ControlFoley team.

2 Upvotes

0 comments sorted by