Skip to main content

Together AI

Paid

Build what's next on the AI Native Cloud.

Together AI is a full-stack AI Native Cloud platform that accelerates inference, model shaping, and pre-training on a research-optimized infrastructure. It provides serverless and dedicated inference for 200+ open-source models spanning text, image, video, audio, and embeddings, alongside GPU clusters, fine-tuning, code sandboxes, and managed storage—all without long-term infrastructure commitments.

inferenceapigpufine tuningopen sourcemultimodalCode GenerationCode RefactoringImage GenerationVideo GenerationText TranslationgenerationeditingtranslationCode AssistantWriting & ContentImage & VisionVideo Creation
Visit Website
Together AI preview

What is it

Together AI is a cloud platform purpose-built for AI-native workloads. It offers serverless inference on a shared fleet of open models, dedicated endpoints on reserved GPUs, on-demand and reserved GPU clusters for training, a managed fine-tuning platform, secure code sandboxes, and high-performance managed storage. The platform is powered by cutting-edge research including FlashAttention-4 and ATLAS runtime-learning accelerators.

What it can do

Developers can run inference on 200+ open-source models through a per-token API with no provisioning or minimum cost. Batch processing handles up to 30 billion tokens asynchronously at 50% lower cost. Dedicated endpoints provide guaranteed performance for steady-traffic applications. GPU clusters scale from self-serve instant clusters to thousands of H100, H200, B200, and GB200 nodes. The fine-tuning platform supports LoRA, full fine-tuning, and DPO on models up to 100B+ parameters. Code sandboxes enable secure execution of LLM-generated code, and managed storage offers zero-egress-fee object storage optimized for AI workloads.

Who is it for

It is designed for AI researchers, ML engineers, AI-native startups, and enterprise product teams who need industrial-grade performance for open-source models without managing their own GPU infrastructure. Teams building applications on Llama, DeepSeek, Qwen, Gemma, and other open-weight models use Together AI for both prototyping and production-scale deployment.

Key Features

Serverless Inference – 200+ open models on demand

Access over 200 open-source models through a shared per-token API with no provisioning, no replicas to size, and no minimum cost. The catalog spans chat (DeepSeek V4 Pro, Qwen3.7-Max, GLM-5.1, Kimi K2.6, Llama 3.3 70B), code (Qwen3-Coder, DeepCoder, GPT-OSS), vision (Gemma 4 31B), image generation (FLUX.2, Imagen 4.0, Stable Diffusion 3), video (Veo 3.0, Kling 2.1, Wan 2.2, Sora 2), speech (Whisper, Orpheus), and embeddings. Cached input discounts apply automatically on supported chat models.

Batch Inference API – Process billions of tokens at half the cost

Submit massive workloads asynchronously through the batch processing API for up to 50% lower cost than real-time serverless rates. Scale to 30 billion tokens per model with any serverless model or private deployment. Ideal for data processing pipelines, large-scale content generation, and overnight analytics where real-time responses are not required.

Dedicated Inference – Guaranteed performance on reserved GPUs

Deploy models on dedicated, single-tenant GPU instances with guaranteed performance, autoscaling, and traffic spike handling. Choose from H100 80GB ($6.49/hr), H200 140GB, or HGX B200 180GB ($11.95/hr). Best for applications with steady traffic, consistent latency requirements, or serving custom fine-tuned models that cannot run on shared infrastructure.

GPU Clusters – Self-serve to thousands of GPUs

Scale from instant on-demand clusters to reserved capacity spanning 7 to 180+ days. Hardware options include NVIDIA HGX H100 ($5.49/hr on-demand, $3.99/hr reserved), HGX H200 ($6.79/hr on-demand, $4.55/hr reserved), HGX B200 ($9.95/hr on-demand, $9.09/hr reserved), GB200 NVL72, and GB300 NVL72. All clusters are optimized with the Together Kernel Collection for better performance than standard cloud GPU instances.

Use Cases

AI Chatbot Deployment

Deploy conversational AI using open-source chat models like DeepSeek V4 Pro, Qwen3.7-Max, or Llama 3.3 70B. Serverless inference handles variable traffic without provisioning, while dedicated endpoints ensure consistent latency for production applications. Cached input discounts reduce costs for multi-turn conversations.

Code Generation and Assistants

Build coding assistants using Qwen3-Coder, DeepCoder, or GPT-OSS models. The Code Interpreter API enables agentic workflows where the AI writes, runs, and debugs code iteratively. Code Sandboxes provide secure development environments for testing generated code at scale.

Image and Video Generation at Scale

Generate marketing visuals, product mockups, and social content using FLUX.2, Imagen 4.0, and Stable Diffusion 3. Create short video clips with Veo 3.0, Kling 2.1, or Wan 2.2. Batch processing reduces costs by 50% for large content pipelines.

Custom Model Fine-Tuning

Fine-tune open models on proprietary data to create domain-specific assistants, style-consistent image generators, or specialized code models. Upload datasets, select hyperparameters, and Together AI handles the training infrastructure. Serve fine-tuned models through the same API without code changes.

Pricing plans

Frequently Asked Questions

Discussion

No comments yet. Be the first to start the thread.