Image-only MiniMax H3: text to image, image editing, up to nine reference images, and native H3 Fun ControlNet. No video endpoint or Colab dependency.
Description excerpted from the original listing, which is linked below.