gen‑ai.news

Open source

Week of 10 August 2026

This week's picks center on reducing inference steps and local fine-tuning friction for video, audio, and image workflows.

  • 1

    larryvrh· 573 likes

    A LoRA adapter for MiniMax-H3 that generates synchronized video and stereo audio together from text prompts, cutting the required sampling steps from roughly 20 down to 4-8. Builders looking to reduce inference costs on joint audio-video generation without a dedicated training run may find the v4 checkpoint a practical starting point, with the caveat that heavy fast-motion scenes at 4 steps benefit from falling back to the older v1-850 weights.

    text-to-videotext-to-audioaudio-videoloraminimax-h3comfyui
  • 2
    unslothGitHub

    unslothai· 69,785 stars

    Unsloth provides a local interface for running and fine-tuning both text and diffusion models, covering options like FLUX for image generation alongside LLMs such as DeepSeek and Gemma. Builders working on diffusion model workflows who want to fine-tune locally without relying on cloud infrastructure may find it a practical starting point.

    agentchatgptdeepseekfine-tuninggemmagemma3
  • 3

    lightx2v· 246 likes

    Minimax-H3-Turbo is a diffusers-compatible image-to-video model from lightx2v, built on top of MiniMaxAI's MiniMax-H3 base and supporting both English and Chinese prompts across text-to-video, image-to-video, and reference-to-video pipelines. Builders already working in the diffusers ecosystem can drop it in with minimal friction and reproduce results directly from the accompanying open repository.

    diffuserst2vi2vr2vimage-to-videoen
  • 4

    op7418· 23,639 stars

    An AI agent skill for generating HTML-based slide decks, covering editorial magazine and Swiss grid layouts, image prompt generation, social covers, and a WebGL presentation runtime. Builders working on agentic pipelines who need structured visual output beyond plain text may find the layout system and low-power runtime a useful reference point.

    ai-agentclaude-codecodexhtml-deckimage-generationppt
  • 5

    drbaph· 240 likes

    A ComfyUI-compatible set of LoRA conversions for MiniMax-H3 Turbo, enabling joint video and synchronized audio generation from text in as few as 8 sampling steps. Builders already using pruned MiniMax-H3 checkpoints in ComfyUI can drop in the recommended v4 step-600 EMA weights with a ready-made workflow, making fast turnaround iteration more accessible without rebuilding a pipeline from scratch.

    minimax-h3loraadaptercomfyuiprunedpruned-model