New

    Introducing AutoEvals: Find the right model for your task

    DeepSeekText Generation Model

    DeepSeek V4 Pro (0813)

    The August 2026 refresh of DeepSeek V4 Pro, served as a distinct pinned version. Same flagship sparse Mixture-of-Experts with a 1M-token context, 131K max output, and low, high, and max reasoning modes for dialing test-time compute per request.

    import OpenAI from "openai";
    
    const openai = new OpenAI({
      baseURL: "https://api.inference.net/v1",
      apiKey: "<YOUR_API_KEY>",
    });
    
    const completion = await openai.chat.completions.create({
      model: "deepseek-v4-pro-0813",
      messages: [
        {
          role: "user",
          content: "What is the meaning of life?"
        }
      ],
      stream: true,
    });
    
    for await (const chunk of completion) {
      process.stdout.write(chunk.choices[0]?.delta.content as string);
    }
    Prompt caching
    1. Keep the prompt prefix stable and byte-identical across calls.
    2. Do not send prompt_cache_key. These routes do not support it and may reject the request.
    3. Check usage.prompt_tokens_details.cached_tokens in the response to confirm cache reads.

    Start building with DeepSeek V4 Pro (0813) today

    Serverless, OpenAI-compatible, and billed per token. No commitments.