New

    Introducing AutoEvals: Find the right model for your task

    GoogleMultimodal Model

    Gemini 3.1 Pro Preview

    Google's Gemini 3 frontier reasoning model, refined for factual grounding and precise tool use. Optimized for software engineering and multi-step agentic workflows over a 1M-token context with text, image, audio, video, and PDF input.

    import OpenAI from "openai";
    
    const openai = new OpenAI({
      baseURL: "https://api.inference.net/v1",
      apiKey: "<YOUR_API_KEY>",
    });
    
    const completion = await openai.chat.completions.create({
      model: "gemini-3.1-pro-preview",
      messages: [
        {
          role: "user",
          content: "What is the meaning of life?"
        }
      ],
      stream: true,
    });
    
    for await (const chunk of completion) {
      process.stdout.write(chunk.choices[0]?.delta.content as string);
    }
    Prompt caching
    1. Keep the stable system content block byte-identical across calls.
    2. Mark it with cache_control: { "type": "ephemeral" } to set a cache point.
    3. Long prompts are required. As a practical target, use roughly 4,000 stable tokens for Gemini 3.1 Pro Preview or 9,000 for the other listed Gemini models; these are not guaranteed provider minimums.
    4. Check usage.prompt_tokens_details.cached_tokens in the response to confirm cache reads.
    {
      "messages": [{
        "role": "system",
        "content": [{
          "type": "text",
          "text": "Stable system content...",
          "cache_control": { "type": "ephemeral" }
        }]
      }]
    }

    Start building with Gemini 3.1 Pro Preview today

    Serverless, OpenAI-compatible, and billed per token. No commitments.