Request#
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.flua.ink/v1",
api_key=os.environ["FLUA_API_KEY"],
)
response = client.chat.completions.create(
model="gpt-6-astra",
messages=[
{"role": "user", "content": "Explain what a token is in one paragraph."},
],
)
print(response.choices[0].message.content)Parameters#
| Parameter | Type | Description |
|---|---|---|
model* | string | Model id from the catalog, e.g. claude-opus-5-5. |
messages* | array | The conversation: objects with role (system, user, assistant, tool) and content. |
max_tokens | integer | Upper bound for the answer length, in tokens. |
reasoning_effort | string | How hard the model thinks: none, low, medium or high. See Reasoning. |
temperature | number | Randomness, usually 0–2. Lower is more precise, higher more varied. |
top_p | number | Nucleus sampling, an alternative to temperature. |
stream | boolean | Stream the answer in chunks over Server-Sent Events. |
stream_options | object | {"include_usage": true} adds token usage to the last stream chunk. |
stop | string | array | Sequences that stop generation. |
tools | array | Functions the model may call. |
tool_choice | string | object | auto, none, required or a specific function. |
response_format | object | JSON mode or a JSON Schema for structured output. |
Response#
{
"id": "chatcmpl-…",
"object": "chat.completion",
"created": 1790964838,
"model": "gpt-6-astra",
"choices": [
{
"index": 0,
"message": { "role": "assistant", "content": "…" },
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 18,
"completion_tokens": 104,
"total_tokens": 122
}
}Streaming#
With stream: true the server sends data: events carrying pieces of the answer in delta.content and ends with data: [DONE]. With include_usage the final chunk carries usage.
data: {"object":"chat.completion.chunk","choices":[{"index":0,"delta":{"role":"assistant"}}]}
data: {"object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"A token"}}]}
data: {"object":"chat.completion.chunk","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
data: {"object":"chat.completion.chunk","choices":[],"usage":{"prompt_tokens":18,"completion_tokens":104,"total_tokens":122}}
data: [DONE]Reasoning#
Reasoning models think before they answer. reasoning_effort sets how much: none turns thinking off where the model allows it, high lets it think longer — more accurate, but slower and pricier. Without the parameter the model decides.
response = client.chat.completions.create(
model="gemini-3.8-flash",
reasoning_effort="high",
messages=[{"role": "user", "content": "What is 17 × 23?"}],
)
message = response.choices[0].message
print(getattr(message, "reasoning_content", None)) # the thinking, if the model returns it
print(message.content)
print(response.usage.completion_tokens_details.reasoning_tokens)- Reasoning tokens are billed as output and reported in
usage.completion_tokens_details.reasoning_tokens. - Gemini, Claude and DeepSeek return the thinking itself in
message.reasoning_content, ordelta.reasoning_contentwhen streaming. GPT models reason privately: only the token count is visible. - Some models always reason (DeepSeek V4.1 Flash, for one) — the parameter barely changes them.
"reasoning": {"effort": "high"}; in the Anthropic format use "thinking": {"type": "enabled", "budget_tokens": 4000}.Tool calling#
Describe functions in tools. When the model decides to call one, the response contains tool_calls — run it and send the result back as a message with role: tool.
import json, os
from openai import OpenAI
client = OpenAI(base_url="https://api.flua.ink/v1", api_key=os.environ["FLUA_API_KEY"])
tools = [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Current weather in a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
},
}]
messages = [{"role": "user", "content": "What's the weather in Kyiv?"}]
first = client.chat.completions.create(model="claude-sonnet-5-5", messages=messages, tools=tools)
call = first.choices[0].message.tool_calls[0]
messages.append(first.choices[0].message)
messages.append({
"role": "tool",
"tool_call_id": call.id,
"content": json.dumps({"temperature_c": 14, "sky": "cloudy"}),
})
final = client.chat.completions.create(model="claude-sonnet-5-5", messages=messages, tools=tools)
print(final.choices[0].message.content)Image input#
Multimodal models accept images as part of content — by URL or as a base64 data URL.
response = client.chat.completions.create(
model="gemini-3.8-flash",
messages=[{
"role": "user",
"content": [
{"type": "text", "text": "What is in this image?"},
{"type": "image_url", "image_url": {"url": "https://example.com/photo.jpg"}},
],
}],
)Structured output#
To get output that follows a schema, pass response_format with type json_schema (for models that support it).
response = client.chat.completions.create(
model="gpt-6-astra",
messages=[{"role": "user", "content": "Invent a name and a tagline for a coffee shop"}],
response_format={
"type": "json_schema",
"json_schema": {
"name": "brand",
"schema": {
"type": "object",
"properties": {"name": {"type": "string"}, "tagline": {"type": "string"}},
"required": ["name", "tagline"],
},
},
},
)