Alle 210+ KI-Modelle
Jedes Bild-, Video-, Audio-, Dokument- und 3D-Modell, das bei faktry verfügbar ist - an einem Ort.

Elevenlabs Speech To Text Scribe V2
ElevenLabs Scribe v2 - Advanced speech-to-text with speaker diarization

Elevenlabs TTS Eleven V3
ElevenLabs TTS v3 - High quality voices

Minimax Music 3
MiniMax Music 3 - Generate a song from lyrics and a music description
Google Gemini Omni Flash
Gemini Omni Flash v1.1 model for text-to-video generation with in-prompt audio control, up to 4K resolution
Google Gemini Omni Flash
Gemini Omni Flash v1.1 model for animating a single reference image into video, with optional end-frame interpolation, up to 4K resolution
Google Gemini Omni Flash Edit
Gemini Omni Flash v1.1 model for instruction-based video editing (e.g. "Make this video anime. Keep everything else the same."), up to 4K resolution

Topaz Generative
Topaz Starlight diffusion-based video upscaling that regenerates lost detail — Precise/HQ/Mini/Sharp for restoration, Fast 2 for cheaper turnaround

Kling Video Ai Avatar V2 Standard
Standard AI Avatar V2 — animate a portrait image with audio using Kling
Lightricks LTX 2.5 Audio To Video Pro
LTX 2.5 Pro model for audio-driven video generation — animates a starting image to match the audio, or generates from prompt alone

Kling Video V3 Standard Motion Control
Kling Video v3 Standard model for motion control in video generation
Nano Banana 2
Nano Banana 2 (Gemini 3.1 Flash) model for fast and efficient text-to-image generation
Nano Banana 2 Edit
Nano Banana 2 model for fast and efficient image editing

GPT 5.6 Terra
OpenAI GPT-5.6 Terra model for fast and efficient image analysis

Claude Sonnet 5
Anthropic's Claude Sonnet 5 model for reverse-engineering image prompts

Pixelcut Background Removal
Pixelcut's fast, ultra high-quality background remover, ideal for e-commerce and image editing workflows

Topaz Generative
Middle tier - Topaz Generative rebuilds missing detail, with Wonder 3/3.5 realism, Redefine prompt-guided detail, and Recovery for extreme low-resolution sources. Wonder 3 and 3.5 bill at half the rate of the other variants.

Claude Sonnet 5
Anthropic's Claude Sonnet 5 model for balanced performance and cost in text generation

Meshy V6
Meshy 6 – realistic production-ready 3D from text

Meshy V7
Meshy 7 – latest-generation 3D from a single image, with optional rigging and animation