Documentation · Chat Completions

API

Chat Completions

The main endpoint for text models, in the OpenAI format: conversations, streaming, tools and image input.

POST/v1/chat/completions

Request#

import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.flua.ink/v1",
    api_key=os.environ["FLUA_API_KEY"],
)

response = client.chat.completions.create(
    model="gpt-6-astra",
    messages=[
        {"role": "user", "content": "Explain what a token is in one paragraph."},
    ],
)
print(response.choices[0].message.content)

Parameters#

ParameterTypeDescription
model*stringModel id from the catalog, e.g. claude-opus-5-5.
messages*arrayThe conversation: objects with role (system, user, assistant, tool) and content.
max_tokensintegerUpper bound for the answer length, in tokens.
reasoning_effortstringHow hard the model thinks: none, low, medium or high. See Reasoning.
temperaturenumberRandomness, usually 0–2. Lower is more precise, higher more varied.
top_pnumberNucleus sampling, an alternative to temperature.
streambooleanStream the answer in chunks over Server-Sent Events.
stream_optionsobject{"include_usage": true} adds token usage to the last stream chunk.
stopstring | arraySequences that stop generation.
toolsarrayFunctions the model may call.
tool_choicestring | objectauto, none, required or a specific function.
response_formatobjectJSON mode or a JSON Schema for structured output.
Supported parameters depend on the model. Parameters a model does not support may be ignored.

Response#

200 OK
{
  "id": "chatcmpl-…",
  "object": "chat.completion",
  "created": 1790964838,
  "model": "gpt-6-astra",
  "choices": [
    {
      "index": 0,
      "message": { "role": "assistant", "content": "…" },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 18,
    "completion_tokens": 104,
    "total_tokens": 122
  }
}

Streaming#

With stream: true the server sends data: events carrying pieces of the answer in delta.content and ends with data: [DONE]. With include_usage the final chunk carries usage.

text/event-stream
data: {"object":"chat.completion.chunk","choices":[{"index":0,"delta":{"role":"assistant"}}]}

data: {"object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"A token"}}]}

data: {"object":"chat.completion.chunk","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}

data: {"object":"chat.completion.chunk","choices":[],"usage":{"prompt_tokens":18,"completion_tokens":104,"total_tokens":122}}

data: [DONE]

Reasoning#

Reasoning models think before they answer. reasoning_effort sets how much: none turns thinking off where the model allows it, high lets it think longer — more accurate, but slower and pricier. Without the parameter the model decides.

reasoning.py
response = client.chat.completions.create(
    model="gemini-3.8-flash",
    reasoning_effort="high",
    messages=[{"role": "user", "content": "What is 17 × 23?"}],
)

message = response.choices[0].message
print(getattr(message, "reasoning_content", None))  # the thinking, if the model returns it
print(message.content)
print(response.usage.completion_tokens_details.reasoning_tokens)
  • Reasoning tokens are billed as output and reported in usage.completion_tokens_details.reasoning_tokens.
  • Gemini, Claude and DeepSeek return the thinking itself in message.reasoning_content, or delta.reasoning_content when streaming. GPT models reason privately: only the token count is visible.
  • Some models always reason (DeepSeek V4.1 Flash, for one) — the parameter barely changes them.
In the Responses format use "reasoning": {"effort": "high"}; in the Anthropic format use "thinking": {"type": "enabled", "budget_tokens": 4000}.

Tool calling#

Describe functions in tools. When the model decides to call one, the response contains tool_calls — run it and send the result back as a message with role: tool.

tools.py
import json, os
from openai import OpenAI

client = OpenAI(base_url="https://api.flua.ink/v1", api_key=os.environ["FLUA_API_KEY"])

tools = [{
    "type": "function",
    "function": {
        "name": "get_weather",
        "description": "Current weather in a city",
        "parameters": {
            "type": "object",
            "properties": {"city": {"type": "string"}},
            "required": ["city"],
        },
    },
}]

messages = [{"role": "user", "content": "What's the weather in Kyiv?"}]
first = client.chat.completions.create(model="claude-sonnet-5-5", messages=messages, tools=tools)
call = first.choices[0].message.tool_calls[0]

messages.append(first.choices[0].message)
messages.append({
    "role": "tool",
    "tool_call_id": call.id,
    "content": json.dumps({"temperature_c": 14, "sky": "cloudy"}),
})
final = client.chat.completions.create(model="claude-sonnet-5-5", messages=messages, tools=tools)
print(final.choices[0].message.content)

Image input#

Multimodal models accept images as part of content — by URL or as a base64 data URL.

vision.py
response = client.chat.completions.create(
    model="gemini-3.8-flash",
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "What is in this image?"},
            {"type": "image_url", "image_url": {"url": "https://example.com/photo.jpg"}},
        ],
    }],
)

Structured output#

To get output that follows a schema, pass response_format with type json_schema (for models that support it).

json_schema.py
response = client.chat.completions.create(
    model="gpt-6-astra",
    messages=[{"role": "user", "content": "Invent a name and a tagline for a coffee shop"}],
    response_format={
        "type": "json_schema",
        "json_schema": {
            "name": "brand",
            "schema": {
                "type": "object",
                "properties": {"name": {"type": "string"}, "tagline": {"type": "string"}},
                "required": ["name", "tagline"],
            },
        },
    },
)