New

    Introducing AutoEvals: Find the right model for your task

    Z.aiReasoning Model

    GLM-5.3 Flash

    Z.ai's GLM-5.3 Flash: a fast open-weight reasoning model with image input, a 1M-token context, and 131K max output. Reasoning effort is selectable from none through max, with prompt caching for high-volume agentic and coding workloads.

    CONTACT

    Meet with our research team

    Schedule a call with our research team to learn more about how Specialized Language Models can cut costs and improve performance.