Core¶
Runtime¶
The one loop, and the policies that make it the lead or a role.
helioai.runtime.runner ¶
The agent loop, written once.
Runner.run is the turn-taking skeleton both stream_chat and stream_subagent used
to carry as separate copies: call the model with a compacted history, start the turn's
tool calls together, dispatch each in the model's order, review the figures, emit the
events, append the results, retry once when the answer quotes ids that exist in no
catalogue, stop at the cap. What differs between the lead and a role is a Policy and
two small hooks — how a tool call may be intercepted before it reaches the registry
(the lead's task and internal tools, a role's whitelist), and what to do with each
model call's usage.
The run yields the same {"event", "data"} dicts the loops yielded, in the same order,
and ends with a RunEnd the wrapper turns into its own closing events — the lead's
reply … done, a role's sub_agent_end.
RunEnd
dataclass
¶
How a run ended, handed to the wrapper as the last item of Runner.run.
Attributes:
| Name | Type | Description |
|---|---|---|
final_text |
str | None
|
The model's closing text; |
turns |
int
|
Model calls made. |
artifacts |
list[dict]
|
Every artifact the run produced, without a role's |
usage |
dict
|
Token counts summed over the run's model calls, as the providers reported them. |
capped |
bool
|
The turn budget ran out before the model answered. |
empty |
bool
|
The model returned neither text nor a tool call and the policy stops there (the lead's output-budget failure). |
claims |
list[dict]
|
The numbers the model says its answer states, each with |
Source code in helioai/runtime/runner.py
Runner
dataclass
¶
One run of the loop under one policy.
Attributes:
| Name | Type | Description |
|---|---|---|
policy |
Policy
|
What makes this run the lead or a role. |
llm |
LLMClient
|
The provider client the run calls. |
intercept |
Intercept | None
|
Called with each tool call and the turn before the registry is
consulted. Returns |
on_llm_call |
OnLLMCall | None
|
Called with the turn and the model's reply after each call; the
lead records the usage row there. The run also sums usage into |
registry |
ToolRegistry
|
Where tool calls are dispatched. The process-wide registry unless a caller hands in another — the wrappers pass their own module's reference, which is what the tests that stub a registry patch. |
ctx |
RunContext | None
|
Who runs, in which session, writing where. Bound to the workspace contextvars for the whole run — the one place they are set — and the source of the directories the writing tools receive as trusted arguments. None leaves the ambient bindings alone, for a caller that manages its own. |
artifacts, |
(usage, turns)
|
Progress so far, readable while the run is in flight and after it raised — a role that blew up still reports what it measured. |
Source code in helioai/runtime/runner.py
207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 | |
run
async
¶
Drive the model over history until it answers, stops or hits the cap.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
history
|
list[Message]
|
The conversation so far, appended to in place — the assistant's replies and every tool result land here, so the caller persists the same list it passed. |
required |
Yields:
| Type | Description |
|---|---|
AsyncIterator[dict | RunEnd]
|
Event dicts, then one |
Source code in helioai/runtime/runner.py
visible_tools ¶
The definitions the model is shown this call: the policy's tools minus the
deferred ones it has not asked for, plus search_tools while any are withheld.
Source code in helioai/runtime/runner.py
search_loop_correction ¶
The note handed to a role that keeps searching instead of downloading.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
searches
|
int
|
How many lookups it has made. |
required |
ids
|
list[str]
|
The product ids its searches already returned, best first. |
required |
Returns:
| Type | Description |
|---|---|
str
|
The correction text; the ids are listed so the next turn has nothing to look up. |
Source code in helioai/runtime/runner.py
helioai.runtime.policies ¶
What makes the one loop the lead agent or a delegated role.
A Policy is data: the prompt, the tools the model sees and the ones it may actually
call, the turn budget, the first turn's tool_choice, and the few behavioural switches
the two loops disagreed on. The lead's policy is built from the constants of
agent_loop; a role's from sub_agents.AGENT_ROLES, which stays the one place roles are
declared.
Policy
dataclass
¶
How one run of the loop behaves.
Attributes:
| Name | Type | Description |
|---|---|---|
name |
str
|
|
system_prompt |
str
|
The instructions placed before the history on every call. |
tools |
tuple[ToolDef, ...]
|
The definitions the model is shown. |
max_turns |
int
|
How many model calls a run may make before it is capped. |
tool_choice_first |
str
|
|
allowed |
frozenset[str] | None
|
The tools a run may call. |
sandbox_no_network |
bool
|
Whether |
sub_agent_ctx |
Mapping[str, str] | None
|
|
comment_replies |
bool
|
Whether text the model writes alongside tool calls is shown as
a |
stream_replies |
bool
|
Whether the model's text is shown as it is generated
( |
stop_on_empty_reply |
bool
|
Whether a response with neither text nor tool calls ends the run as a failure. The lead's does — the usual cause is the output budget, and the error names the setting to raise; a role's simply ends with an empty summary. |
bogus_retry |
bool
|
Whether an answer quoting ids absent from the catalogue buys the model one more turn, with the correction appended, before it is accepted. |
provider |
str | None
|
The provider this run's client talks to — for the usage rows, which bill per provider. None means the lead's configured one. |
model |
str | None
|
The model, when the run was given one of its own ( |
deferred |
frozenset[str]
|
Tools whose definitions are withheld from the model until it asks for
them with |
final_answer |
bool
|
Whether the model may close the run with |
search_budget |
int
|
How many lookups ( |
search_tools_names |
frozenset[str]
|
The tools that count as lookups. |
data_tools_names |
frozenset[str]
|
The tools whose first call ends the lookup phase. |
Source code in helioai/runtime/policies.py
event_extra
property
¶
The keys added to every event of the run: a role's context, or nothing.
helioai.runtime.context ¶
Who is running, in which session, writing where — as an object, not an ambience.
Three contextvars (workspace._current_user, _current_session, _current_label) used
to be set by each loop, by the MCP server and by nobody in a test, and read back from
inside the tools, the datastore and the sandbox: an argument nobody passed and everybody
depended on. A tool called from a test wrote under the default user; a sub-agent worked
only because the lead had bound the label first; a save landed in a temporary directory
when the caller forgot the session.
RunContext carries the same facts explicitly and is handed to the runner, which binds
them for the duration of a run — bound() is the one place the three contextvars are
set and reset. The tools that write receive their directories as trusted arguments
(tool_exec.trusted_args); the contextvars remain the hot path for everything else,
RunContext.current() reads them back, and the fallback logs when it is used.
RunContext
dataclass
¶
The facts a run writes under.
Attributes:
| Name | Type | Description |
|---|---|---|
user_id |
str
|
Owner of every path derived during the run. |
session_id |
str
|
The conversation; a sub-agent runs under its parent's. |
session_dir |
Path
|
The workspace directory — figures, scripts, npz, manifest, ledger.
A sub-agent shares its lead's, which is how |
label |
str | None
|
The human-readable directory name, when the session has one. |
agent |
str
|
|
task_id |
str | None
|
The delegation's correlation id, for a sub-agent. |
no_network |
bool
|
Whether |
query |
str | None
|
The user's question for this turn, verbatim. It used to stop at the lead's
loop: a sub-agent saw only the lead's brief, and nothing downstream could ask
whether a result answered what the person typed. Carried on the context, it
reaches every sub-agent through |
Source code in helioai/runtime/context.py
28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 | |
catalogs_dir
property
¶
Where the user's saved catalogues live: beside the workspaces, not in one.
for_session
classmethod
¶
Build a context from ids, deriving the session directory the way
workspace.get_session_dir does: by label when there is one, else by id.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
user_id
|
str
|
Owner of the session. |
required |
session_id
|
str
|
The conversation. |
required |
label
|
str | None
|
Its workspace directory name, when already minted. |
None
|
**extra
|
Any
|
|
{}
|
Source code in helioai/runtime/context.py
current
classmethod
¶
The context the ambient contextvars describe, or None when no session is bound.
The transition path: a caller that has not been handed a context yet can still obtain the one its caller bound.
Source code in helioai/runtime/context.py
child ¶
A sub-agent's context: the same user, session and directory, its own identity.
Source code in helioai/runtime/context.py
bound ¶
Bind the workspace contextvars to this context for the duration of a block.
Tokens are reset in reverse order, so a sub-agent bound inside its lead's run hands the lead's bindings back untouched.
Source code in helioai/runtime/context.py
helioai.runtime.validator ¶
One verdict on a finished answer: its ids, its recipes, its figures, its numbers.
Four checks judged the lead's answer from four places — _flag_unknown_ids and
_flag_recipe_bypass in tool_exec, provenance_check.check_reply for the numbers in
the prose, vision.maybe_review for the figures — each with its own event and none aware
of the others. validate runs them in one place and returns a Verdict the wrapper turns
into the events the interfaces already render, plus one verdict event that carries the
whole judgement.
What is new is the judgement of claims: when the model closed with final_answer, it
named each number and where it came from. Those are compared to the provenance ledger by
name, with a unit-aware tolerance — 2.5 states a recorded 2.519, 57.2 deg states
57.16 deg, 0.0025 pT states 2.5 nT — instead of being found again in the prose and
attributed by the words around them. The regex check stays as the net under the prose:
a number the model stated without claiming it is still judged.
Verdict
dataclass
¶
How an answer holds up, in every respect the runtime can check.
Attributes:
| Name | Type | Description |
|---|---|---|
matched |
list[dict]
|
Claims a ledger entry of the same name states (unit-aware). |
contradicted |
list[dict]
|
Claims naming a recorded scalar that holds another value. |
unsourced |
list[dict]
|
Claims nothing in the session computed — including those the model
itself marked |
unknown_ids |
list[str]
|
Parameter ids the answer quotes that exist in no catalogue. |
recipe_flags |
list[dict]
|
Exports that look like a calibrated recipe's output while the recipe was never loaded, or loaded and never called. |
figure_reviews |
list[str]
|
The vision verdicts on the run's figures. |
prose |
dict | None
|
The regex-based provenance report on the free text, or None when the session computed nothing or the text states no number. |
Source code in helioai/runtime/validator.py
as_event ¶
The verdict event payload: counts up front, every detail behind them.
Source code in helioai/runtime/validator.py
validate ¶
validate(text: str, claims: list[dict], *, history: list, artifacts: list[dict], session_dir: Path, figure_reviews: list[str] | None = None) -> tuple[str, Verdict]
Judge a finished answer.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
text
|
str
|
The answer, as the model wrote it. |
required |
claims
|
list[dict]
|
The numbers it named ( |
required |
history
|
list
|
The turn's messages, read for evidence a recipe was loaded. |
required |
artifacts
|
list[dict]
|
What the run exported, to tell a computed value from a quoted one. |
required |
session_dir
|
Path
|
Whose ledger to check against. |
required |
figure_reviews
|
list[str] | None
|
The vision verdicts collected during the run. |
None
|
Returns:
| Type | Description |
|---|---|
str
|
The text — annotated when a check appends a correction, exactly as the two |
Verdict
|
checks always annotated it — and the verdict. |
Source code in helioai/runtime/validator.py
judge_claim ¶
Place one claim against the ledger: matched, contradicted or unsourced.
The entries considered are those whose name is the claim's source or name — the
export it says it came from. Any run that produced the value sources it (a session
that exported compression_ratio twice, 3.045 then 2.538, computed both), so the
entries are read newest first and the first that states the value wins, within the
prose checker's tolerance (RTOL) or the rounding the claim shows. A named scalar
that states another value contradicts the claim; an entry holding several values
cannot accuse, since a reply legitimately quotes one component of a vector; nor can
a dimensioned scalar accuse a claim that gives no units — the claim may be another
quantity of the same run, and on the first live run it was (a normal's components
filed under the angle's export). A claim the model marked literature or
asserted is never contradicted: it did not say the session computed it.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
claim
|
dict
|
|
required |
entries
|
list[dict]
|
|
required |
Returns:
| Type | Description |
|---|---|
tuple[str, dict]
|
The status and the detail dict shown to a reader. |
Source code in helioai/runtime/validator.py
helioai.runtime.plan ¶
The plan the model announced, kept as data, and how the run then followed it.
present_plan(title, steps) was a display: the loop forwarded it to the interfaces as a
plan event and forgot it. What the model then did was left to the reader to compare
with what it had said — three tool calls later nobody remembers step 2 named
get_timeseries, and a plan that promised the theta_bn recipe and ran hand-written
arithmetic instead looked exactly like one that was followed. Plan is the announced
plan as data; adherence places the turn's tool calls against it and reports, in one
plan_report event, which planned tools were used, which were not, and which tools were
used without being planned. The report describes; it never blocks, corrects or retries.
Two things the first live run settled. A step's tool field is prose — the model wrote
"search_parameters + get_timeseries" and "task (data_analyst)" — so it is read word by
word against the tools the lead knows, not compared whole. And a planned tool the lead
had a sub-agent run is a step done, not a deviation: the lead planned the analysis in its
analyst's tools and then delegated every step, and the report said 0/4 used, unplanned:
task. A sub-agent's calls therefore satisfy the plan; they are never unplanned — the
second run planned in delegations alone (task (data_analyst)) and its analyst's six
tools were all "unplanned" — because the task step covers whatever the role does, and
the role's whitelist, not the lead's plan, governs it. What the lead does with its own
hands is held to the plan, task excepted: delegating is how a step gets done. The
scaffolding calls (the plan itself, the skills, search_tools, final_answer) count for
nothing on either side.
Step
dataclass
¶
One announced step.
Attributes:
| Name | Type | Description |
|---|---|---|
description |
str
|
What the step does, as the model put it. |
tool |
str | None
|
The tool it said it would use, or None when it named none. |
Source code in helioai/runtime/plan.py
Plan
dataclass
¶
A plan as present_plan delivered it, with the payload's looseness removed.
Attributes:
| Name | Type | Description |
|---|---|---|
title |
str
|
The plan's one-line title. |
steps |
tuple[Step, ...]
|
The steps in order; a step the model wrote as a bare string is kept as a description without a tool. |
Source code in helioai/runtime/plan.py
from_payload
classmethod
¶
Build a plan from a present_plan payload or a plan event's data.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
data
|
dict
|
|
required |
Returns:
| Type | Description |
|---|---|
Plan
|
The plan, with malformed steps dropped rather than raised on. |
Source code in helioai/runtime/plan.py
tools ¶
The tools the plan names, once each, in the order of their first mention.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
known
|
Collection[str] | None
|
The tools the lead can call. A step's |
None
|
Returns:
| Type | Description |
|---|---|
list[str]
|
Tool names, first mention first. |
Source code in helioai/runtime/plan.py
calls_made ¶
The tools a turn called: the lead's own, and its sub-agents', each once in order.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
events
|
Iterable[dict]
|
The turn's events; |
required |
Returns:
| Type | Description |
|---|---|
tuple[list[str], list[str]]
|
|
Source code in helioai/runtime/plan.py
delegations_made ¶
How each delegation of the turn ended: the role, its turns, whether it was capped.
A plan written in delegations alone ("task (data_analyst)") is always followed by delegating, so the tools say nothing; what a reader wants to know is whether the role finished. The run after the Runner extraction is the case: the librarian hit its four-turn cap, was re-delegated and finished in two — visible in the trace, invisible in a count of tools.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
events
|
Iterable[dict]
|
The turn's events; the lead's own |
required |
Returns:
| Type | Description |
|---|---|
list[dict]
|
|
Source code in helioai/runtime/plan.py
adherence ¶
Compare what the run did with what the plan said.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
plan
|
Plan
|
The plan the turn opened with. |
required |
events
|
Iterable[dict]
|
The turn's events, as yielded. |
required |
known
|
Collection[str] | None
|
The tools the lead can call, to read the plan's |
None
|
Returns:
| Type | Description |
|---|---|
dict
|
The |
dict
|
|
dict
|
sub-agents called), |
dict
|
|
dict
|
excepted), |
dict
|
of planned tools that were called, or None when the plan named no tool and |
dict
|
there is nothing to hold the run to. |
Source code in helioai/runtime/plan.py
Agent loop¶
helioai.core.agent_loop ¶
The agent decision loop.
Given a user message and a session id, run the LLM in a tool-using loop: the LLM may emit tool calls, execute them via the ToolRegistry, feed the results back, and iterate until the LLM produces a final text reply (or we hit the safety cap).
Two consumption modes share the same generator core (stream_chat): - chat() → collects all events, returns a single ChatResult - stream_chat() → async generator, yields one event dict per step
Event kinds and their payloads are listed once, in core/events.py, and held to the
emitters and the three renderers by tests/test_events_contract.py.
ChatResult
dataclass
¶
Final outcome of a non-streaming chat() call.
Source code in helioai/core/agent_loop.py
build_lead_system_prompt ¶
Return the lead agent system prompt.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
restricted
|
bool
|
True (the public default) appends the scope guardrail, so the model refuses off-topic requests itself. False is reached only with a valid dev token and yields the base prompt. |
required |
experiments
|
frozenset[str] | None
|
The experiments in force; |
None
|
Returns:
| Type | Description |
|---|---|
str
|
The full system prompt text. |
Source code in helioai/core/agent_loop.py
stream_chat
async
¶
stream_chat(llm_client: LLMClient, user_id: str, session_id: str, user_text: str, *, restricted: bool = True) -> AsyncIterator[dict]
Run one conversational turn of the agent and stream its progress as events.
This is the package's central API: every interface (CLI, web SSE, Jupyter, MCP) is a consumer of this generator. History is loaded from and persisted to the session store keyed by (user_id, session_id), so consecutive calls with the same ids continue the same conversation.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
llm_client
|
LLMClient
|
Provider client from |
required |
user_id
|
str
|
Storage namespace — workspaces and profiles live under it. |
required |
session_id
|
str
|
Conversation id; reuse it to continue, mint one to start fresh. |
required |
user_text
|
str
|
The user's message for this turn. |
required |
restricted
|
bool
|
True (default) appends the heliophysics scope guardrail; False (dev token) exposes the base prompt only. |
True
|
Yields:
| Type | Description |
|---|---|
AsyncIterator[dict]
|
Dicts with an |
AsyncIterator[dict]
|
|
AsyncIterator[dict]
|
appended to the session's journal ( |
AsyncIterator[dict]
|
yielded, so |
Example
llm = build_llm_client() async for ev in stream_chat(llm, "cli", "my-session", "IMF Bz at L1 today?"): ... if ev["event"] == "reply": ... print(ev["text"], end="")
Source code in helioai/core/agent_loop.py
chat
async
¶
chat(llm_client: LLMClient, user_id: str, session_id: str, user_text: str, *, restricted: bool = True) -> ChatResult
Run one agent turn to completion and return the final result.
Non-streaming wrapper over stream_chat — same arguments, same session
semantics — for callers that want the answer, not the progress feed
(Jupyter magic, scripts, tests).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
llm_client
|
LLMClient
|
Provider client from |
required |
user_id
|
str
|
Storage owner; decides which tree under |
required |
session_id
|
str
|
Conversation to append to. An unknown id starts a new one. |
required |
user_text
|
str
|
The question. |
required |
restricted
|
bool
|
Whether the scope guardrail is in force. See
|
True
|
Returns:
| Type | Description |
|---|---|
ChatResult
|
ChatResult with |
ChatResult
|
|
Example
result = await chat(build_llm_client(), "cli", "my-session", ... "Plasma beta for B=5 nT, n=10 cm^-3, T=20 eV?") print(result.reply) # final assistant text (model-dependent) len(result.artifacts) # figures / parameter cards / code produced
Source code in helioai/core/agent_loop.py
Sub-agents¶
helioai.core.sub_agents ¶
Sub-agents for delegating focused heliophysics subtasks.
Each role declares a tool whitelist, a system addon, and an optional set of skills auto-loaded into the sub's system prompt. Sub-agents run in isolation with a fresh context — the lead's history is invisible to them.
The events a sub-agent yields are the non-lead-only kinds of core/events.py.
SubAgentRole
dataclass
¶
A specialised agent: its prompt, its tool whitelist and its turn budget.
allowed_tools is enforced, not advisory — a role calling outside its set
gets an error naming what it may use, and the tool is never dispatched.
search_budget is the role's lookup allowance before it must touch data; it only
takes effect under the search_budget experiment, and is 0 (no limit) otherwise.
Source code in helioai/core/sub_agents.py
task_tool_def ¶
Build the task tool definition offered to the lead agent.
Deliberately not registered in the ToolRegistry: the agent loop intercepts
task and spawns a sub-agent instead of dispatching a function.
Returns:
| Type | Description |
|---|---|
ToolDef
|
A ToolDef whose |
Source code in helioai/core/sub_agents.py
stream_subagent
async
¶
stream_subagent(role: str, description: str, *, parent_session_id: str, user_id: str, llm_client: LLMClient, task_id: str | None = None, context: RunContext | None = None) -> AsyncIterator[dict]
Async generator that runs a sub-agent and yields progress events.
Yields the kinds of core/events.py that are not lead-only (tool_call, tool_result,
skill_loaded, artifact, figure_review, invalid_ids, recipe_bypassed), each enriched
with sub_agent_ctx={role, task_id}, then a final sub_agent_end carrying
findings/summary/artifacts/n_iterations/error.
summary is the sub-agent's whole deliverable and is emitted in full: it becomes
the lead's tool result, so anything cut here is a measurement the lead can no
longer report and will be tempted to invent. Callers that display it truncate on
their own side. Reaching the turn cap is reported as an error, not as a result,
for the same reason — a lead handed a capped run has nothing to summarise.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
role
|
str
|
One of the configured roles. Its tool whitelist and turn cap are enforced, not advisory. |
required |
description
|
str
|
The task, in the lead's own words. |
required |
parent_session_id
|
str
|
The lead's session. The sub-agent writes into that
same workspace, which is how |
required |
user_id
|
str
|
Storage owner, inherited from the lead. |
required |
llm_client
|
LLMClient
|
Provider client, shared with the lead. |
required |
task_id
|
str | None
|
Correlation id echoed in every event, so a caller running several sub-agents can tell their streams apart. |
None
|
context
|
RunContext | None
|
The lead's run context; the sub-agent runs under a child of it, in the same session directory. None derives one from the ids and the bound label, for callers that predate contexts. |
None
|
Yields:
| Type | Description |
|---|---|
AsyncIterator[dict]
|
Progress events, then a final |
Source code in helioai/core/sub_agents.py
330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 | |
Tool execution¶
The tool-call mechanics the runner uses.
helioai.core.tool_exec ¶
Shared tool-execution helpers for the agent loops.
Both the lead loop (agent_loop.stream_chat) and the sub-agent loop (sub_agents.stream_subagent) run the same tool-call mechanics: inject the sandbox run dir for run_python, summarise the result, detect skill loads, and extract renderable artifacts. Keeping that logic here — imported by both loops — prevents the two copies from drifting apart (a real bug source, see the session-13 _extract_artifact list/dict regression).
This module imports neither agent_loop nor sub_agents, so there is no cycle.
compact_history ¶
Return a copy of messages where tool-result messages older than the last
keep_full are summarized. The most recent results stay verbatim (the next LLM
call usually needs them); older ones — already consumed — are trimmed so context
does not grow unbounded over a long session. Persisted history is left untouched;
only the per-call payload shrinks.
A loaded recipe is never summarized. It is the one result the model keeps writing
code against for the rest of the task, and two run_python calls later it had been
cut to its first line: in two live runs the analyst reloaded it, and once rewrote the
formula from memory rather than call the function it could no longer see. A recipe
is a few kilobytes; keeping it costs less than the extra turn.
The cap per stale result is 1 500 characters, up from 300. Measured on the live runs of 00_quickstart: every tool result of a four-turn session together is 14 kB, less than the system prompt and tool schemas re-sent on every call, and the numbers are exempt from the cap anyway. What 300 bought was a lead that could not recall the previous cell's method and an analyst that had lost its own exports.
ponytail: fixed window N=2; widen keep_full if a case regresses on stale results.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
messages
|
list
|
The per-call payload, newest last. |
required |
keep_full
|
int
|
How many of the most recent tool results to leave verbatim. Zero summarises every one of them, which is the degraded retry a context-length failure would want. |
2
|
Returns:
| Type | Description |
|---|---|
list
|
A new list. |
list
|
persisted. |
Source code in helioai/core/tool_exec.py
trusted_args ¶
The framework-injected arguments of the tools that write to disk.
Passed via call_tool(..., trusted=...), so they bypass the private-argument guard
that rejects model- or MCP-supplied _* overrides. run_python and run_recipe get
their workspace, run index and network flag; get_timeseries and
get_events_timeseries the session's data directory; save_catalog the user's
catalogue directory. Every other tool gets nothing. Reading them off the context
rather than off the workspace contextvars is what lets a tool called from a test, or
over MCP, write where its caller said and nowhere else.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
name
|
str
|
Tool about to be called. |
required |
ctx
|
RunContext | None
|
The run's context. None falls back to the bound contextvars, the way the loops resolved the session before contexts existed. |
None
|
no_network
|
bool
|
Deny the sandbox a network namespace; |
False
|
Returns:
| Type | Description |
|---|---|
dict
|
The trusted argument dict, empty for a tool that writes nothing. |
Source code in helioai/core/tool_exec.py
start_tool_calls ¶
start_tool_calls(tool_calls: list[ToolCall] | None, *, allowed: set[str] | frozenset[str] | None = None, registry: ToolRegistry | None = None, ctx: RunContext | None = None) -> dict[str, asyncio.Task]
Start every parallel-safe registry call of a turn at once, keyed by call id.
The prompt asks the model to batch its downloads in one turn, and the loops then
ran them one after the other: a data_analyst's first turn — three or four
get_timeseries — took the sum of their durations. Started here, they overlap;
the caller still awaits each result in the model's order, so every event and
every tool message keeps the order it had when the calls were sequential.
Skipped, and left to the caller's sequential path: run_python (see
_SEQUENTIAL_TOOLS), anything not in the registry (the task tool, the internal
tools), and — for a sub-agent — anything outside its whitelist, which the caller
refuses without dispatching.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
tool_calls
|
list[ToolCall] | None
|
The assistant's tool calls for this turn. |
required |
allowed
|
set[str] | frozenset[str] | None
|
A sub-agent's whitelist; None for the lead, who may call anything. |
None
|
registry
|
ToolRegistry | None
|
Where the calls are dispatched; the process-wide registry by default. |
None
|
ctx
|
RunContext | None
|
The run's context, from which the data tools get their directory as a
trusted argument — the same |
None
|
Returns:
| Type | Description |
|---|---|
dict[str, Task]
|
|
dict[str, Task]
|
a |
dict[str, Task]
|
result — so awaiting a task is safe. |
Source code in helioai/core/tool_exec.py
cancel_pending ¶
Cancel the calls a turn started and never awaited — a cancelled turn must not leave downloads running for nobody.
Source code in helioai/core/tool_exec.py
emit_post_tool_events ¶
emit_post_tool_events(name: str, result: ToolResult, *, tool_result_extra: dict | None = None, common_extra: dict | None = None) -> Iterator[dict]
Yield the events that follow a completed tool call.
Order is tool_result → (skill_loaded if load_skill) → artifact(s),
matching what both loops emitted before this was factored out.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
name
|
str
|
The tool that just ran. |
required |
result
|
ToolResult
|
Its result; the payload is read for artifacts, the model's text for the summary the event carries. |
required |
tool_result_extra
|
dict | None
|
Merged into the |
None
|
common_extra
|
dict | None
|
Merged into |
None
|
Yields:
| Type | Description |
|---|---|
dict
|
The events, in the order both loops emitted them before this was |
dict
|
factored out: |
Source code in helioai/core/tool_exec.py
unknown_id_correction ¶
The correction handed back when an answer quotes ids that are not in the catalogue.
Shared between the retry that buys the model another turn and the annotation left on an answer that has run out of turns, so both say exactly the same thing.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
bogus
|
list[str]
|
The ids that are absent from the index. |
required |
Returns:
| Type | Description |
|---|---|
str
|
The correction text, without the answer it refers to. |
Source code in helioai/core/tool_exec.py
recipe_available ¶
recipe_available(tool_name: str, arguments: dict | None, result: ToolResult, history: list) -> ToolResult
Tell the model, on the run_python result it is about to read, that a shipped
recipe covers what its code just computed.
_flag_recipe_bypass reads the same signals at the end of the turn and says so in
the answer — after the model has stopped acting, for the reader. Here the same
finding rides on the tool result itself, so the next turn can call run_recipe
instead of defending a hand-written copy: on the fourth live run of 00_quickstart
the analyst loaded theta_bn, rewrote the formula inline, exported theta_bn, and
reported 54.85° from a window the recipe would not have chosen; nothing told it
before its answer. Same constants and helpers, imported, so the two readings cannot
disagree. Annotates, never blocks — the code ran, its exports stand, and the model
may still argue. A run_recipe result is exempt: the shipped source ran.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
tool_name
|
str
|
The tool that just ran. |
required |
arguments
|
dict | None
|
Its arguments; the |
required |
result
|
ToolResult
|
Its result; the exports say what the code computed. |
required |
history
|
list
|
The run so far, for the recipes it loaded or ran. |
required |
Returns:
| Type | Description |
|---|---|
ToolResult
|
The result, with |
ToolResult
|
the payload when a signal fires; unchanged otherwise. |
Source code in helioai/core/tool_exec.py
check_answer ¶
Confront a finished answer with the catalogue and the recipe shelf.
Both loops call this, which is the whole point of it living here. Both checks were
written inside the sub-agent loop and stayed there, so a lead agent that did the
physics itself — Acts III and IV of the showcase notebook, on a run where
load_recipe was called zero times all session — was never checked at all. The
detectors were not silent because the run was clean; they were silent because
nothing called them.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
text
|
str
|
The finished answer. |
required |
history
|
list
|
The turn's messages, read for evidence a recipe was loaded and then bypassed. |
required |
artifacts
|
list[dict]
|
What the run exported, used to tell a computed value from a quoted one. |
required |
Returns:
| Type | Description |
|---|---|
tuple[str, list[str], list[dict]]
|
The text (annotated when a check fires), the unknown ids, and the recipe flags. |
Source code in helioai/core/tool_exec.py
Sessions¶
helioai.core.session ¶
Conversation history keyed by (user_id, session_id), persisted to SQLite.
SessionStore ¶
Conversation history keyed by (user_id, session_id), persisted to SQLite.
Histories are cached in memory per key and written back whole on save.
Tests use a real database on tmp_path rather than a mock: a mocked store
passed happily through a schema migration that broke production.
Example
store = SessionStore(tmp_path / "sessions.db") history = store.get_or_create("cli", "sess-1") # [] on first call history.append(Message(role="user", content="hello")) store.save("cli", "sess-1", history) store.get_or_create("cli", "sess-1")[0].role 'user'
Source code in helioai/core/session.py
105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 | |
turn_lock ¶
The lock a caller must hold for the whole of one conversational turn.
get_or_create hands every caller the same list, and a turn is a
read-modify-write of it that spans several awaits: without this, two turns on
one session — two browser tabs, two MCP calls — interleave their appends and
save persists the mix. _lock only serialises the SQL, never the turn.
One asyncio.Lock per key, created on first use. A lock in Python ≥ 3.10 binds
to an event loop only on its first contended acquisition, so the CLI and the
Jupyter magic — a fresh asyncio.run per question, never two turns on one
session at once — reuse it across loops safely, while the web and MCP servers
run one loop. If that assumption ever breaks, asyncio raises a RuntimeError
naming the loop mismatch instead of silently corrupting anything.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
user_id
|
str
|
Owner of the session. |
required |
session_id
|
str
|
Session identifier. |
required |
Returns:
| Type | Description |
|---|---|
Lock
|
The same lock object for the same key, for the life of the store. |
Source code in helioai/core/session.py
is_busy ¶
Whether a turn is currently running for this session.
Read without taking the lock — this is the web layer's fast refusal (409), not
a guarantee; the guarantee is turn_lock itself.
Source code in helioai/core/session.py
get_or_create ¶
Return the cached history for a session, loading it from disk if needed.
Source code in helioai/core/session.py
save ¶
Replace a session's stored history.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
user_id
|
str
|
Owner of the session. |
required |
session_id
|
str
|
Session identifier. |
required |
history
|
list[Message]
|
Full message list; it replaces whatever was stored. |
required |
Source code in helioai/core/session.py
record_usage ¶
record_usage(user_id: str, session_id: str, *, turn: int | None, agent: str, provider: str, prompt_tokens: int, completion_tokens: int, cached_tokens: int = 0) -> None
Append what one LLM call cost, as the provider reported it.
The counts have ridden on Message for a while and were dropped at save time,
so a session reloaded from disk reported no cost and nothing could say what a
user had spent. Kept apart from messages on purpose: one row per call, never
rewritten by save, so a compacted or reset history does not erase the bill.
Zero counts are skipped — a provider that reports none leaves no row rather
than a misleading zero.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
user_id
|
str
|
Owner of the session. |
required |
session_id
|
str
|
The conversation the call belonged to; a sub-agent's calls are charged to its parent session. |
required |
turn
|
int | None
|
Turn index within the run, for ordering. |
required |
agent
|
str
|
|
required |
provider
|
str
|
Provider name, since a session may switch providers. |
required |
prompt_tokens
|
int
|
Input tokens billed. |
required |
completion_tokens
|
int
|
Output tokens billed. |
required |
cached_tokens
|
int
|
The part of the prompt served from the provider's cache. |
0
|
Source code in helioai/core/session.py
usage_totals ¶
Sum a user's token usage, optionally for one session or a recent window.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
user_id
|
str
|
Whose usage. |
required |
session_id
|
str | None
|
Restrict to one session; None for every session. |
None
|
since_days
|
float | None
|
Only calls in the last N days; None for all time. |
None
|
Returns:
| Type | Description |
|---|---|
dict
|
|
dict
|
when nothing was recorded. |
Source code in helioai/core/session.py
append_event ¶
Journal one event of a turn, in the order it was yielded.
The journal is what a session replays from: the browser used to rebuild a
past conversation by re-parsing the JSON of every tool message with a
hundred lines of shape-sniffing, and lost the plan, the provenance verdict, the
figure reviews and everything a sub-agent did on the way. Written before the
event is handed to the interface, so a stream cut mid-turn still leaves what was
shown. Its own table, never rewritten by save, so a compacted history does
not erase the record.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
user_id
|
str
|
Owner of the session. |
required |
session_id
|
str
|
The conversation the event belongs to. |
required |
event
|
dict
|
|
required |
Source code in helioai/core/session.py
events ¶
A session's journal, oldest first, in the shape the loops yield.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
user_id
|
str
|
Owner of the session. |
required |
session_id
|
str
|
The conversation. |
required |
Returns:
| Type | Description |
|---|---|
list[dict]
|
|
list[dict]
|
turn after turn, so a replay renders with the same code as the live stream. |
list[dict]
|
Empty for a session recorded before the journal existed. |
Source code in helioai/core/session.py
reset ¶
Delete a session, its messages, its usage rows and its journal, and drop it from the cache.
Source code in helioai/core/session.py
set_workspace_dir ¶
Record which workspace directory a session's artifacts live in.
Source code in helioai/core/session.py
get_workspace_dir ¶
Return a session's workspace directory label, or None.
Source code in helioai/core/session.py
workspace_dirs ¶
All workspace dir labels owned by a user (for path-ownership checks).
Source code in helioai/core/session.py
all_sessions ¶
Return a user's session ids, most recently updated first.
Source code in helioai/core/session.py
list_summaries ¶
Summarise a user's recent sessions for the history view.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
user_id
|
str
|
Owner of the sessions. |
required |
limit
|
int
|
Maximum number of sessions to return. |
50
|
Returns:
| Type | Description |
|---|---|
list[dict]
|
Dicts with session_id, updated_at, first_message, n_messages, |
list[dict]
|
workspace_dir and tokens (prompt + completion, all calls), most recent first. |
Source code in helioai/core/session.py
strip_orphan_tool_calls ¶
Remove assistant tool_calls that have no matching tool response.
An interrupted generation (e.g. client disconnect mid-tool) can leave an assistant message with tool_calls but no corresponding tool messages in the history. Sending such a sequence to the LLM API causes a 400 error.
For each orphaned tool_call id: - If the assistant message has content too, keep the message but drop the orphaned tool_calls list entry (or clear it entirely if all are orphaned). - If the assistant message has no content and all its tool_calls are orphaned, drop the message entirely.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
history
|
list[Message]
|
Messages in order, as loaded from the session store. |
required |
Returns:
| Type | Description |
|---|---|
list[Message]
|
A new list; the input is not modified. Messages without tool calls pass |
list[Message]
|
through untouched, so a clean history is returned equal to its input. |
Source code in helioai/core/session.py
Skills¶
helioai.core.skills_loader ¶
Discover and serve markdown-defined skills to the agent.
Each skill lives in skills/
SkillError ¶
SkillMeta
dataclass
¶
Header of a skill, as listed to the agent before it loads the body.
Source code in helioai/core/skills_loader.py
load_index ¶
Return the markdown index of available skills, for the agent to browse.
Source code in helioai/core/skills_loader.py
load_skill ¶
Return a skill's full markdown body.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
name
|
str
|
Skill name as listed by |
required |
Returns:
| Type | Description |
|---|---|
str
|
The skill body, ready to append to a system prompt. |
Raises:
| Type | Description |
|---|---|
SkillError
|
If the skill does not exist or the name escapes the skills directory. |
Source code in helioai/core/skills_loader.py
list_skill_names ¶
list_skills ¶
Judgment¶
The seam for a System One judge beside the loop, and the joins that read its contract. See The judgment layer.
helioai.core.judgment ¶
One seam between HelioAI and a System One judge — every question and every threshold.
Why one file: HelioAI is PyHC-listed and open source, and a reviewer who asks "what does
this agent delegate to a proprietary model?" must be able to answer by reading one module.
Every site that consults the judge builds its questions here-shaped (Noul, Choice),
calls ask, and reads an Answers — or None.
Why abstention is None and not a sentinel: a sentinel can be compared, summed and sorted
into a plausible wrong answer; None > 0.5 raises. The judge abstains when the backend is
null (the default), when the site's experiment is off, when the call fails or times out,
and per question when the answer does not clear that question's own threshold. Every call
site reads if answers is None: <today's code> — so a build without a key, without the
extra, or with the default backend is bit-for-bit the loop that exists today. That is the
property the tests hold this module to.
Why observation only: the 2026-09-15 lesson — nothing that changes what the model sees
ships without a replay on both sides of the same question — applies to every answer here.
So ask records one JSON line per call under the session workspace (judgment.jsonl),
with the state in full: a disagreement that cannot be adjudicated later without the state
is not a measurement. What a site does with an answer is the site's business, and until a
site has earned it, that is: annotate, never correct.
Why the model call is bounded: a judgment that can hold a turn is a judgment that can
hang it. The whole call sits under settings.judgment.timeout_s (round trip measured
from France on 2026-09-22: median 266 ms, p95 373 ms), and a timeout abstains.
DELIVERABLES
module-attribute
¶
What a request may want delivered, each asked as its own yes/no: one Choice over the same set abstained on every request that wanted two things at once — "plot |B| and compute θ_Bn" is a figure and a value — and abstaining on the commonest shape of a question is not a contract. A request that wants none of the four is an explanation.
Noul
dataclass
¶
A yes/no question.
The judge returns a probability; the answer is True at or above 0.5 + margin,
False at or below 0.5 - margin, and None in between. The default margin makes
0.7 the least confident "yes" — a site that needs more certainty raises it.
Attributes:
| Name | Type | Description |
|---|---|---|
instructions |
str
|
The question, in plain English, as the judge reads it. |
margin |
float
|
Half-width of the abstention band around 0.5. |
Source code in helioai/core/judgment.py
Choice
dataclass
¶
One option out of a closed set.
The judge returns the chosen option and a confidence; below floor the answer is
None. The options are the vocabulary the site already speaks — for a measurement
type, the index's own measurement_type values — so an answer is usable as an exact
filter, never as prose to interpret.
Attributes:
| Name | Type | Description |
|---|---|---|
instructions |
str
|
The question, in plain English. |
options |
tuple[str, ...]
|
The closed set, in the order the site wants them presented. |
floor |
float
|
Minimum confidence for the choice to count as an answer. |
Source code in helioai/core/judgment.py
Answers
dataclass
¶
What the judge said for one call, already decided question by question.
values holds bool for a Noul, str for a Choice, None where the answer did
not clear its threshold; raw keeps the probabilities so the record can be re-judged
at another threshold later without another call.
Attributes:
| Name | Type | Description |
|---|---|---|
values |
Mapping[str, bool | str | None]
|
Decided answers, keyed like the questions. |
raw |
Mapping[str, Any]
|
The judge's probabilities, keyed like the questions. |
model |
str
|
The judge model that answered. |
latency_ms |
float
|
Wall time of the call as this process saw it. |
request_id |
str | None
|
The provider's id for the call, for support and audit. |
Source code in helioai/core/judgment.py
get ¶
The decided answer for name, or default when absent or abstained.
enabled ¶
Whether site may ask at all: a judging backend AND the site's experiment name.
Two axes on purpose. With the backend alone, "Jev does not help at this site" and "the layer costs something" could not be told apart; with the experiment alone, a site could be on with nobody to answer.
Source code in helioai/core/judgment.py
ask
async
¶
ask(site: str, state: Mapping[str, Any], questions: Mapping[str, Question], *, decided: Any = None) -> Answers | None
Ask the judge; None means abstain, and the caller runs today's code.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
site
|
str
|
The call site's name; |
required |
state
|
Mapping[str, Any]
|
What the judge reads — JSON-serialisable, and recorded in full. |
required |
questions
|
Mapping[str, Question]
|
The questions, keyed by the names the site will read back. |
required |
decided
|
Any
|
What the deterministic path decided, when the site knows it before asking; recorded beside the answer so the two can be compared offline. |
None
|
Returns:
| Type | Description |
|---|---|
Answers | None
|
The decided answers, or |
Source code in helioai/core/judgment.py
batch
async
¶
batch(site: str, states: list[Mapping[str, Any]], questions: Mapping[str, Question], *, concurrency: int = 8, record_to: Path | None = None, keys: Sequence[str] | None = None) -> list[Answers | None]
Ask the same questions of many states, off the agent loop — for a job, not a turn.
The runtime's ask is gated by a site's experiment name because it runs inside a
conversation nobody asked to be judged. A job such as helioai index --classify is an
explicit request: it needs only a judging backend, and it records every call to a file
of its own (record_to) rather than to a session that does not exist. The same
thresholds decide the answers, so what a job writes and what a turn would read agree.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
site
|
str
|
The job's name, in the records. |
required |
states
|
list[Mapping[str, Any]]
|
One state per item, JSON-serialisable. |
required |
questions
|
Mapping[str, Question]
|
The questions, asked of every state. |
required |
concurrency
|
int
|
How many calls in flight at once (37 ms per item at eight, measured). |
8
|
record_to
|
Path | None
|
The JSON-lines file every call is appended to; None records nothing. |
None
|
keys
|
Sequence[str] | None
|
One name per state, written into its record as |
None
|
Returns:
| Type | Description |
|---|---|
list[Answers | None]
|
One |
Source code in helioai/core/judgment.py
intent_contract
async
¶
The contract for one question, or None — see INTENT_QUESTIONS.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
query
|
str
|
The user's question, verbatim. |
required |
Source code in helioai/core/judgment.py
contract_fields ¶
The intent event's payload: the decided answers, the date assembled by the code.
A Choice of none, or an abstention, is None in the payload — the reader must not
mistake "the judge did not say" for "the request named nothing". date is the ISO day,
month or year the components make, with date_precision saying which; deliverable
is the kinds wanted joined with + ("value+figure"), explanation when all four were
decided no, None when any of them was left undecided and none was yes.
Source code in helioai/core/judgment.py
collect
async
¶
Read a contract task the turn started; abstain if it has not answered in time.
The task ran concurrently with the model; by the time the answer is out it has almost always finished. A turn never waits on the judge longer than one judgment budget.
Source code in helioai/core/judgment.py
to_sdk ¶
Our question shapes as typesafe_sdk ones — the only place the SDK's names appear.
Options of a Choice become criteria without descriptions: the vocabulary is the
site's own (an index field, a closed list of regions), and describing it in prose
would put a second, unmeasured question in front of the judge.
Source code in helioai/core/judgment.py
aclose
async
¶
Close the judge's connection pool at shutdown; safe when it was never opened.
Source code in helioai/core/judgment.py
helioai.core.joins ¶
What the turn did about what was asked — four joins, no model, observation only.
The intent contract (judgment.intent_contract) says what a question committed the answer
to. This module places that contract against what the turn actually produced — the
parameter cards of every product loaded, the claims of final_answer, the figures — and
records the comparison on the intent event. Each join is an exact operation on typed
fields: a set membership, an interval intersection, a count. None calls a model, none
reads prose, and none changes the answer; the intent event is what a reader of the
journal, or a later correction path behind its own experiment, will consult.
The four joins are the failures the benches actually produced, one each:
- frame — the frame the question named against the
coord_sysof the cards the sandbox filled from the archive's metadata: "GSM asked,BGSEplotted" was caught by nothing. - window — the date the question named against the bounds the downloads obtained,
and the cards whose series stopped short of the window asked for (
coverage_note). The 2026-09-18 θ_Bn of 12° for a 54° shock had a window that ended at the shock; every check downstream was green because none looked at the bounds. - quantity — the measurement type the question named against the indexed type of the products loaded: twelve searches and a confident answer built on no product that measured the quantity (2026-09-15) is a count, once the field is filled.
- responsiveness — what the question required against what the answer delivered: an uncertainty asked and no claim about a spread, a figure asked and none produced.
A join that has nothing to compare — no frame named, no card with a bound, no type on
any loaded product — reports None, never a verdict: "the join could not say" and "the
turn got it right" are kept apart, as they are in the contract itself.
checks ¶
The four joins for one turn, as the checks field of the intent event.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
contract
|
dict
|
|
required |
artifacts
|
list[dict]
|
|
required |
claims
|
list[dict]
|
|
required |
Source code in helioai/core/joins.py
summary ¶
The mismatches, as short phrases for a renderer; empty when every join agreed or had nothing to compare.
Source code in helioai/core/joins.py
Figure review¶
helioai.core.vision ¶
Stateless vision side-call: review sandbox figures after run_python.
The image is sent once, outside the conversation; only the short text verdict enters the tool result (and thus the history), so the cost stays a few hundred tokens per figure instead of re-sending images every turn. Never blocks the loop: any failure logs a warning and returns the result unchanged.
maybe_review
async
¶
Attach a vision verdict to a run_python result carrying figures.
No-op unless HELIOAI_VISION_ENABLED is set and the tool is run_python.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
tool_name
|
str
|
Name of the tool that just ran. |
required |
result
|
ToolResult
|
Its result; the payload is read for |
required |
Returns:
| Type | Description |
|---|---|
ToolResult
|
(possibly amended result, verdict text or None) — the verdict is a |
str | None
|
stateless side-call; only its text enters the history, never the image. |
Source code in helioai/core/vision.py
LLM clients¶
helioai.core.llm.base ¶
Provider-neutral message model and LLMClient interface.
ToolCall
dataclass
¶
A tool invocation requested by the model.
Attributes:
| Name | Type | Description |
|---|---|---|
id |
str
|
Provider-assigned identifier, echoed back on the matching tool
result. Gemini has no native ids, so its client synthesises
|
name |
str
|
Registered tool name. |
arguments |
dict
|
Decoded JSON arguments. Empty when the model emitted malformed JSON — a bad tool call must not kill the loop. |
Source code in helioai/core/llm/base.py
Message
dataclass
¶
One turn of conversation, in a provider-neutral form.
Every client converts to and from this shape, so the agent loop, the session store and the interfaces never see a provider's wire format.
Attributes:
| Name | Type | Description |
|---|---|---|
role |
Literal['system', 'user', 'assistant', 'tool']
|
Who produced the turn. |
content |
str
|
Text content. Present alongside |
tool_calls |
list[ToolCall] | None
|
Tools the assistant wants invoked, when it requested any. |
tool_call_id |
str | None
|
For |
prompt_tokens |
int
|
Input tokens the provider billed for this reply, 0 when it reported none. |
completion_tokens |
int
|
Output tokens the provider billed for this reply. |
cached_tokens |
int
|
The subset of |
origin |
str | None
|
Who really wrote a |
name |
str | None
|
For |
reasoning |
str | None
|
For |
Source code in helioai/core/llm/base.py
ToolDef
dataclass
¶
A tool as advertised to the model.
Attributes:
| Name | Type | Description |
|---|---|---|
name |
str
|
Tool name the model will call. |
description |
str
|
What the tool does and when to reach for it — the model's only clue about applicability. |
parameters |
dict
|
JSON Schema object describing the accepted arguments. |
Source code in helioai/core/llm/base.py
LLMClient ¶
Bases: ABC
Interface every provider client implements.
One method is required, deliberately: the agent loop only ever needs a single completion with optional tool calling. Streaming happens at the loop level.
Source code in helioai/core/llm/base.py
aclose
async
¶
Release the underlying HTTP connection pool.
Callers that build a client per request — the CLI, the Jupyter magic and
the web endpoints all do — must await this before their event loop ends.
An async pool binds to the loop that used it, so a client left to the
garbage collector schedules its own teardown after asyncio.run has
closed that loop, and asyncio reports an unretrieved
RuntimeError: Event loop is closed while the sockets stay open.
The default is a no-op so a client without a pool needs no override.
Source code in helioai/core/llm/base.py
chat
abstractmethod
async
¶
chat(messages: list[Message], tools: list[ToolDef], system_prompt: str | None = None, tool_choice: str = 'auto') -> Message
Send one turn and return the assistant's reply.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
messages
|
list[Message]
|
Conversation history. |
required |
tools
|
list[ToolDef]
|
Tools the model may call this turn. |
required |
system_prompt
|
str | None
|
Instructions placed before the history. |
None
|
tool_choice
|
str
|
|
'auto'
|
Returns:
| Type | Description |
|---|---|
Message
|
The assistant reply, carrying |
Source code in helioai/core/llm/base.py
stream_chat
async
¶
stream_chat(messages: list[Message], tools: list[ToolDef], system_prompt: str | None = None, tool_choice: str = 'auto') -> AsyncIterator[str | Message]
Send one turn and yield the reply's text as it is generated, then the reply.
The default is chat() in one piece: a client with no streaming support yields
the finished Message and nothing before it, so a caller that streams works
unchanged against every provider. Clients that stream yield text deltas (str)
and end with the same Message chat() would have returned — same content,
same tool calls, same usage.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
messages
|
list[Message]
|
Conversation history. |
required |
tools
|
list[ToolDef]
|
Tools the model may call this turn. |
required |
system_prompt
|
str | None
|
Instructions placed before the history. |
None
|
tool_choice
|
str
|
As for |
'auto'
|
Yields:
| Type | Description |
|---|---|
AsyncIterator[str | Message]
|
Text deltas, then the final |
Source code in helioai/core/llm/base.py
call_with_retry
async
¶
call_with_retry(fn: Callable[[], Awaitable[Any]], *, attempts: int = 4, base_delay: float = 1.0, max_delay: float = 60.0) -> Any
Call async fn with backoff on retryable HTTP errors.
A server-supplied Retry-After wins over the exponential backoff: waiting the window a rate limiter asks for costs less than losing a session that is already several minutes and several downloads deep.
But only when the wait is one we would actually sit through. A per-minute limiter
asks for seconds; an exhausted daily quota answers Retry-After: 35464 — nearly ten
hours. Capping that to max_delay and retrying anyway just buys four minutes of
silence before the same failure, so a window longer than max_delay fails at once
and says how long it really is.
Non-retryable errors and exhausted attempts are re-raised immediately.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
fn
|
Callable[[], Awaitable[Any]]
|
Zero-argument coroutine function performing one attempt. Taking a callable rather than a coroutine is what lets it be re-invoked. |
required |
attempts
|
int
|
Total tries, including the first. |
4
|
base_delay
|
float
|
First backoff, doubled per attempt. |
1.0
|
max_delay
|
float
|
Backoff ceiling, and the threshold above which a server's
own |
60.0
|
Returns:
| Type | Description |
|---|---|
Any
|
Whatever |
Raises:
| Type | Description |
|---|---|
Exception
|
The last error, re-raised once attempts run out or when the error is not retryable. |
Source code in helioai/core/llm/base.py
close_sdk_client
async
¶
Close an SDK client's connection pool, whether its close() is sync or async.
openai.AsyncOpenAI.close is a coroutine; google.genai.Client.close is not.
Failures are swallowed: this only ever runs while tearing down, and a pool
that will not close is not worth crashing a finished analysis over.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
client
|
Any
|
Any SDK client. One exposing neither |
required |
Source code in helioai/core/llm/base.py
helioai.core.llm.openai_compat ¶
Single client for every provider that speaks the OpenAI chat-completions wire format.
Groq, Ollama, Azure OpenAI and OpenAI itself all accept the same request shape, so
they share one implementation here instead of one near-identical class each. A
provider is a base_url plus a couple of dialect flags, not a subclass.
Azure is the one exception that still needs its own SDK client object (deployment
routing and api-version), so it subclasses this to swap the constructor only.
OpenAICompatClient ¶
Bases: LLMClient
Chat client for any OpenAI-compatible endpoint.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
model
|
str
|
Model name sent in the request (the deployment name on Azure). |
required |
api_key
|
str
|
Provider API key. Local endpoints such as Ollama ignore it, but the SDK requires a non-empty value. |
''
|
base_url
|
str | None
|
Endpoint root. |
None
|
system_role
|
str
|
Role used for the system prompt — |
'system'
|
max_output_tokens
|
int
|
Cap on generated tokens. |
4096
|
temperature
|
float | None
|
Sampling temperature. |
0.2
|
provider
|
str
|
Name used to label log messages. |
'openai'
|
client
|
Any
|
Pre-built SDK client. Injected by tests and by subclasses. |
None
|
Example
client = OpenAICompatClient( ... model="llama-3.3-70b-versatile", ... api_key="gsk_...", ... base_url="https://api.groq.com/openai/v1", ... ) reply = await client.chat([Message(role="user", content="hi")], tools=[])
Source code in helioai/core/llm/openai_compat.py
231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 | |
aclose
async
¶
chat
async
¶
chat(messages: list[Message], tools: list[ToolDef], system_prompt: str | None = None, tool_choice: str = 'auto') -> Message
Send one chat turn and return the assistant's reply.
Streamed underneath, with the deltas discarded. The OpenCode gateway drops
reasoning_content from every non-streamed DeepSeek reply (0 of 24 turns on
2026-09-29, 24 of 24 streamed), and DeepSeek then rejects a later request that
does not send it back — so the sub-agents, which do not stream to anyone, died
where the lead did not. Streaming is the one shape every provider here already
serves to the lead, so asking it of every call costs nothing.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
messages
|
list[Message]
|
Conversation history. |
required |
tools
|
list[ToolDef]
|
Tools the model may call. |
required |
system_prompt
|
str | None
|
Instructions prepended as the first message. |
None
|
tool_choice
|
str
|
|
'auto'
|
Returns:
| Type | Description |
|---|---|
Message
|
The assistant reply, carrying |
Source code in helioai/core/llm/openai_compat.py
stream_chat
async
¶
stream_chat(messages: list[Message], tools: list[ToolDef], system_prompt: str | None = None, tool_choice: str = 'auto') -> AsyncIterator[str | Message]
Send one chat turn, yielding the reply's text as it arrives, then the reply.
Text deltas are yielded as the provider sends them, with an inline
<think> block held back until it closes, and a separate reasoning_content
kept whole on Message.reasoning, never yielded. Tool-call fragments are
reassembled by index and parsed as a finished call is; the usage comes from
the final chunk (stream_options.include_usage). An endpoint that rejects the
streaming request is asked again without it. A turn that ends with neither
text nor a tool call is retried once, streamed again: on a reasoning model the
whole allowance can go into hidden reasoning, and the identical request,
replayed, came back with two tool calls — the loop above treats an empty turn
as fatal, so without the retry the question was abandoned.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
messages
|
list[Message]
|
Conversation history. |
required |
tools
|
list[ToolDef]
|
Tools the model may call. |
required |
system_prompt
|
str | None
|
Instructions prepended as the first message. |
None
|
tool_choice
|
str
|
|
'auto'
|
Yields:
| Type | Description |
|---|---|
AsyncIterator[str | Message]
|
Text deltas, then the final |
Source code in helioai/core/llm/openai_compat.py
391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 | |
to_openai_messages ¶
Convert neutral messages to the OpenAI wire format.
System messages already in the history are dropped: the system prompt is
passed separately by chat() so it always lands first.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
messages
|
list[Message]
|
Conversation history in HelioAI's provider-neutral form. |
required |
Returns:
| Type | Description |
|---|---|
list[dict]
|
Message dicts ready to send as the |
Source code in helioai/core/llm/openai_compat.py
to_openai_tools ¶
Convert tool definitions to OpenAI function-calling schemas.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
tools
|
list[ToolDef]
|
Tools the agent may call this turn. |
required |
Returns:
| Type | Description |
|---|---|
list[dict]
|
Function schemas ready to send as the |
Source code in helioai/core/llm/openai_compat.py
from_openai_response ¶
Convert an OpenAI chat-completions response to a neutral message.
Text content is preserved even when tool calls are present, and a tool call
whose arguments are not valid JSON degrades to {} with a warning rather
than raising — a malformed model output must not kill the agent loop. An
inline <think>...</think> reasoning block, when a provider emits one, is
stripped from the content before it reaches the agent loop or the user; a
separate reasoning_content field is kept on Message.reasoning, because
DeepSeek requires it back on every later request that carries tools.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
response
|
Any
|
The SDK response object. |
required |
provider
|
str
|
Provider name, used only to label log messages. |
'openai'
|
Returns:
| Type | Description |
|---|---|
Message
|
The assistant's reply, with |
Source code in helioai/core/llm/openai_compat.py
helioai.core.llm.azure_openai ¶
Azure-hosted OpenAI.
Same wire format as every other OpenAI-compatible provider, so the conversion
logic lives in openai_compat. Azure only differs in how the client is built —
the endpoint embeds the deployment name and an api-version — plus two dialect
details the base class already exposes as parameters: the system prompt is sent
with the developer role, and temperature is omitted when unset because GPT-5
and the o-series reject it.
AzureOpenAIClient ¶
Bases: OpenAICompatClient
Chat client for an Azure OpenAI deployment.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
api_key
|
str
|
Azure OpenAI API key. |
required |
endpoint
|
str
|
Resource root, e.g. |
required |
api_version
|
str
|
Azure API version, e.g. |
required |
deployment
|
str
|
Deployment name, sent as the request's |
required |
max_output_tokens
|
int
|
Cap on generated tokens. |
4096
|
temperature
|
float | None
|
Sampling temperature; |
None
|
Source code in helioai/core/llm/azure_openai.py
helioai.core.llm.gemini ¶
GeminiClient: Google genai SDK → neutral Message model.
GeminiClient ¶
Bases: LLMClient
Chat client for Google Gemini, using the native google-genai SDK.
Kept separate from OpenAICompatClient because the wire format genuinely
differs: turns are Content objects with typed parts, the assistant role is
called model, and tool results are matched by function name rather than
by id. Since Gemini issues no call ids, this client synthesises name::hex
and parses the name back out on the way in — a format that is persisted in
existing sessions, so it must stay readable.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
api_key
|
str
|
Gemini API key. |
required |
model
|
str
|
Model name, e.g. |
required |
max_output_tokens
|
int
|
Cap on generated tokens. |
4096
|
temperature
|
float
|
Sampling temperature. |
0.2
|
Source code in helioai/core/llm/gemini.py
17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 | |
aclose
async
¶
chat
async
¶
chat(messages: list[Message], tools: list[ToolDef], system_prompt: str | None = None, tool_choice: str = 'auto') -> Message
Send one turn to Gemini and return the assistant's reply.
When system_prompt is not given, it is recovered from any system message
left in the history. tool_choice="required" maps to Gemini's ANY mode.
Source code in helioai/core/llm/gemini.py
helioai.core.llm.factory ¶
Build the configured LLM client.
Providers that speak the OpenAI wire format are table entries, not classes — see
OPENAI_COMPAT. Azure and Gemini need their own SDK client objects and stay
explicit below.
build_llm_client ¶
Return a client for the requested provider.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
provider
|
str | None
|
Provider name. Defaults to |
None
|
model
|
str | None
|
Model (or, on Azure, deployment) to use instead of the provider's
configured one — how a delegated role runs on a smaller model than the lead
( |
None
|
Returns:
| Type | Description |
|---|---|
LLMClient
|
A ready-to-use client. |
Raises:
| Type | Description |
|---|---|
RuntimeError
|
If the provider is unknown or its API key is missing. |
Example
llm = build_llm_client("groq") type(llm).name 'OpenAICompatClient'
Source code in helioai/core/llm/factory.py
Storage¶
helioai.datastore ¶
Session datastore — persist downloaded timeseries as .npz for reuse in run_python.
/data/.npz — compressed numpy archive
Datasets are accessible in the sandbox via load_data("name"). The persisted data is the full-resolution download (before any downsampling). All I/O errors are silently swallowed — persistence must never break a tool call.
fill_mask ¶
Boolean mask of samples that carry no measurement.
Three conventions have to be caught at once, which is why every caller shares this one function instead of applying its own threshold:
- non-finite (NaN/inf) — already unusable;
- the ~1e31 magnitude convention, used by ACE among others;
- the value the dataset declares in its CDF FILLVAL, which is the only way to catch Wind/SWE's 99999.9. A blanket "reject >= 99999" rule is wrong: OMNI carries a real proton temperature of 99093 K in the 2003 Halloween window.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
fillval
|
float | list | None
|
the declared FILLVAL, when the provider exposes it. A list as often
as a scalar — Wind/SWE declares |
None
|
Example
fill_mask(np.array([1.2, -1e31, 3.4]), [-1e31]).tolist() [False, True, False]
Source code in helioai/datastore.py
blank_fill ¶
Return (values with fill blanked to NaN, mask), or (values, None) if not numeric.
Applied at every point where downloaded data is persisted, so that anything reading it back — the sandbox, the exported notebook — sees NaN for "no measurement" rather than a sentinel that looks like a plausible reading. Leaving it to the reader meant remembering to call clean(), which cannot see FILLVAL, so one forgotten call put a 99999.9 "speed" into a plot and a mean.
Non-numeric parameters (string labels, epochs) have no fill convention and are passed through untouched.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
values
|
Any
|
Array as the provider delivered it. |
required |
fillval
|
Any
|
Sentinel(s) the archive declares, read from |
None
|
Returns:
| Type | Description |
|---|---|
Any
|
|
Any | None
|
or |
Example
vals, mask = blank_fill(np.array([1.2, -1e31, 3.4]), [-1e31]) vals.tolist(), mask.tolist() ([1.2, nan, 3.4], [False, True, False])
Source code in helioai/datastore.py
dir_lock ¶
The lock to hold across a read-modify-write of the files under path.
A threading.Lock, not an asyncio.Lock: these writers are synchronous, called
from coroutines today and from worker threads tomorrow, and a thread lock is
correct in both places without binding to any event loop. One lock per
directory string, created on first use and never dropped — a few dozen entries
per process at most.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
path
|
Path
|
The directory whose index files the caller is about to rewrite. |
required |
Returns:
| Type | Description |
|---|---|
Lock
|
The same lock for the same directory, for the life of the process. |
Source code in helioai/datastore.py
write_json_atomic ¶
Write payload as JSON so that path is never observed half-written.
The text goes to a sibling .tmp file first and is swapped in with os.replace,
atomic on POSIX and on Windows when the target exists. A crash mid-write used to
leave a truncated manifest.json and with it an unreadable session.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
path
|
Path
|
Destination file. |
required |
payload
|
Any
|
Anything |
required |
Source code in helioai/datastore.py
read_manifest ¶
Return the manifest dict for a given session directory.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
session_dir
|
Path
|
A session workspace directory (contains |
required |
Returns:
| Type | Description |
|---|---|
dict
|
{"datasets": {name: {kind, param_id, start, stop, ...}}} — empty |
dict
|
datasets dict when no manifest exists. |
Source code in helioai/datastore.py
find_existing ¶
Name of an already-persisted dataset for this exact param and window, or None.
The same matching _unique_name uses to reuse a slot — but consulted before the
download rather than after, so a repeat request costs a dict lookup instead of a
network round-trip. The prompt has always said "download each parameter ONCE"; a
real run still re-fetched the same Wind field three times across three turns, with
the dataset name sitting in plain sight in its own history. Discipline the model
does not reliably apply belongs in the tool.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
param_id
|
str
|
Full speasy id, e.g. |
required |
start
|
str
|
ISO start of the window, matched exactly. |
required |
stop
|
str
|
ISO stop of the window, matched exactly. |
required |
data_dir
|
Path | None
|
The session's data directory; None resolves the bound session. |
None
|
Returns:
| Type | Description |
|---|---|
str | None
|
The dataset name to load, or None. The match is exact: an overlapping |
str | None
|
but different window is a miss, because returning a shorter series than |
str | None
|
asked for would be silently wrong. |
Source code in helioai/datastore.py
save_timeseries ¶
save_timeseries(name_hint: str, *, time, values, param_id: str, units: str, start: str, stop: str, columns, source: str, data_dir: Path | None = None) -> dict | None
Persist a timeseries download as npz + a manifest entry.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
name_hint
|
str
|
Basis for the dataset name (slugged, collision-suffixed). |
required |
time
|
array - like
|
Time axis as returned by speasy. |
required |
values
|
ndarray
|
Data array (fill already blanked). |
required |
param_id
|
str
|
Speasy id, recorded in the manifest for the export rewrite. |
required |
units
|
str
|
Physical units string. |
required |
start
|
str
|
ISO window start, recorded for the export rewrite. |
required |
stop
|
str
|
ISO window stop. |
required |
columns
|
list[str]
|
Component names. |
required |
source
|
str
|
Which tool produced the download. |
required |
data_dir
|
Path | None
|
The session's data directory, as the run's context names it; None resolves the bound session and says so in the log. |
None
|
Returns:
| Type | Description |
|---|---|
dict | None
|
{"dataset": |
dict | None
|
persisting failed (the download result is still usable in-memory). |
Source code in helioai/datastore.py
270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 | |
save_event_collection ¶
save_event_collection(name_hint: str, *, series: list[tuple], param_id: str, units: str, source: str, data_dir: Path | None = None) -> dict | None
Persist a batch of per-event timeseries under one dataset name.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
name_hint
|
str
|
Preferred dataset name; a suffix is added if it is taken. |
required |
series
|
list[tuple]
|
|
required |
param_id
|
str
|
The parameter every event was sampled from. |
required |
units
|
str
|
Units as the provider reports them. |
required |
source
|
str
|
Provenance string carried into the manifest. |
required |
data_dir
|
Path | None
|
The session's data directory; None resolves the bound session. |
None
|
Returns:
| Type | Description |
|---|---|
dict | None
|
|
Source code in helioai/datastore.py
357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 | |
helioai.workspace ¶
Workspace — stable output directory for sandbox figures and data.
Figures go to workspace/
The session label is a human-readable slug derived from the first user message, propagated via a contextvar set by stream_chat at the start of each request.
current_user ¶
Return the user owning the current context, or the default user.
Returns:
| Type | Description |
|---|---|
str
|
The bound user id, or |
str
|
and the Jupyter magic never bind one, so unowned callers still get a |
str
|
real storage home rather than an error. |
Source code in helioai/workspace.py
set_user ¶
Bind the user that owns storage for the current context.
A contextvar rather than a global: every task spawned from here inherits the binding, which is what keeps a sub-agent writing into the same user's tree as the lead that spawned it, while a concurrent web request keeps its own.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
user_id
|
str
|
Owner of every path derived until the binding is reset. |
required |
Returns:
| Type | Description |
|---|---|
object
|
An opaque token to hand back to |
object
|
than by re-setting a previous value is what makes nesting safe. |
Source code in helioai/workspace.py
reset_user ¶
Restore the user binding that was in force before set_user.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
token
|
object
|
The value |
required |
Source code in helioai/workspace.py
user_home ¶
A user's private storage home: <data>/users/<user>/.
The directory is not created — callers that only need to read must not materialise a home for a user that does not exist.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
user
|
str
|
Owner id. Not sanitised here; callers that accept it from a
request pass it through |
required |
Returns:
| Type | Description |
|---|---|
Path
|
The path, whether or not anything exists at it. |
Example
user_home("cli") PosixPath('.../data/users/cli')
Source code in helioai/workspace.py
set_session ¶
Bind the session whose workspace the current context writes into.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
session_id
|
str
|
Session id, used as a directory name when no label is bound. |
required |
Returns:
| Type | Description |
|---|---|
object
|
An opaque token for |
Source code in helioai/workspace.py
reset_session ¶
Restore the session binding in force before set_session.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
token
|
object
|
The value |
required |
set_label ¶
Bind the human-readable folder name for the current session.
Takes precedence over the session id in get_session_dir, so a workspace is
findable by what was asked rather than by a uuid.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
label
|
str
|
Directory name, sanitised on use rather than here. |
required |
Returns:
| Type | Description |
|---|---|
object
|
An opaque token for |
Source code in helioai/workspace.py
reset_label ¶
Restore the label binding in force before set_label.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
token
|
object
|
The value |
required |
safe_id ¶
Reduce an identifier to something that cannot escape its parent directory.
Session ids are caller-supplied — a web request body, an MCP client, a CLI
flag — and end up as path components here, in the export filename, and in the
rmtree behind DELETE /api/sessions/{id}. Everything the project mints is a
uuid4, so stripping to [A-Za-z0-9_-] is lossless in practice and turns
../.. into the fallback rather than a parent directory.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
value
|
str
|
Caller-supplied identifier, trusted for nothing. |
required |
fallback
|
str
|
Returned when nothing survives the filter, so the result is never the empty string — which would resolve to the parent directory. |
'session'
|
Returns:
| Type | Description |
|---|---|
str
|
At most 64 characters of |
Example
safe_id("../../etc/passwd"), safe_id("sess-abc-123456") ('etcpasswd', 'sess-abc-123456')
Source code in helioai/workspace.py
make_session_label ¶
Build a human-readable slug for the session workspace folder.
The session id suffix is what keeps two identical questions from sharing one
directory; the words are only there so a human can find it. Six characters of the
id are enough when ids are random, and not when a client chooses them: the web API
accepts any ^[A-Za-z0-9_-]{1,64}$, so session-001 and session-002 — or a
benchmark's bench-<question>-<hex> — collapsed onto one directory, and the second
session found the first one's downloads in its inventory. taken is the set of
labels the user already owns; when the six-character label is among them the suffix
grows until it is not, and falls back to a digest of the whole id.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
first_message
|
str
|
The question that opened the session. |
required |
session_id
|
str
|
Session id; its first six characters are the usual discriminator. |
required |
taken
|
Collection[str]
|
Labels already assigned to this user's other sessions. |
()
|
Returns:
| Type | Description |
|---|---|
str
|
A slug of at most four words plus the id suffix, distinct from every |
Example
make_session_label("Plot IMF Bz from ACE", "abc123def456") 'plot-imf-bz-from_abc123' make_session_label("Plot IMF Bz from ACE", "abc123def456", {"plot-imf-bz-from_abc123"}) 'plot-imf-bz-from_abc123de'
Source code in helioai/workspace.py
session_dir_for ¶
The workspace directory of a session, from its ids rather than from the ambience.
The same rule get_session_dir applies to the bound contextvars — the label when
the session has one, the id otherwise — so a RunContext built from ids and a
caller reading the contextvars land in the same directory.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
user
|
str
|
Owner of the session. |
required |
session_id
|
str
|
The conversation. |
required |
label
|
str | None
|
Its human-readable directory name, when already minted. |
None
|
Returns:
| Type | Description |
|---|---|
Path
|
The directory, created if missing. |
Source code in helioai/workspace.py
get_session_dir ¶
Return the workspace directory for the current session.
Prefers the bound label, falls back to the bound session id, and lands in a
temporary directory when neither is bound — an unbound caller still gets a
writable place rather than an exception, because run_python must not fail
for want of a session.
Returns:
| Type | Description |
|---|---|
Path
|
The directory, created if missing. |
Source code in helioai/workspace.py
get_next_run_idx ¶
Return the next available run index for a session directory.
Derived from what is on disk rather than from a counter, so the numbering survives a restart and stays right when a session is resumed.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
session_dir
|
Path
|
Directory holding the |
required |
Returns:
| Type | Description |
|---|---|
int
|
|
Source code in helioai/workspace.py
get_run_dir_for_sandbox ¶
The current session directory as a string, for the sandbox fallback.
Returns:
| Type | Description |
|---|---|
str
|
|
str
|
argument list, which takes no |
Source code in helioai/workspace.py
is_under_workspace ¶
True if path is safely under the per-user storage root (no traversal).
is_relative_to rather than a string prefix: comparing against str(root) + "/"
hard-coded the POSIX separator, so on Windows the check never matched and /figure
and /code returned 404 for every legitimate path. Fail-closed, so it was a dead
web UI rather than a hole — but dead all the same.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
path
|
str | Path
|
Candidate path, resolved before comparison so symlinks and |
required |
Returns:
| Type | Description |
|---|---|
bool
|
True when the resolved path sits under the users root. False on any |
bool
|
resolution error, which keeps an unreadable path from being served. |
Source code in helioai/workspace.py
cleanup_old_runs ¶
Purge session directories older than the TTL, for every user.
Called on CLI startup, when the notebook magic loads and when the MCP server
starts, and every hour by cleanup_periodically under the web server — the one
process that never restarts, and so never reached this until it did. Removal
failures are ignored rather than raised: housekeeping must not stop a user from
asking a question. A user's home (profile, catalogs, speasy seed) is never touched;
only what is under workspace/.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
ttl_seconds
|
int | None
|
Age above which a session directory is deleted, measured on
its mtime. Defaults to |
None
|
Returns:
| Type | Description |
|---|---|
int
|
How many directories were removed. |
Source code in helioai/workspace.py
cleanup_periodically
async
¶
Run cleanup_old_runs every period_seconds until the task is cancelled.
The web server used to clean up once, at startup, and then run for weeks: a demo machine had 76 session directories of 269 MB each, every one of them past the TTL. The sweep runs in a worker thread so a slow disk never stalls a stream, and a failing sweep is logged and retried at the next tick rather than ending the task.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
period_seconds
|
float
|
Time between sweeps; the first sweep happens after one period, since the caller already swept at startup. |
3600.0
|
Source code in helioai/workspace.py
Export¶
helioai.export ¶
Export a session as a reproducible Jupyter notebook.
A research result that cannot be re-run is worthless. The agent already saves
every sandbox run as code_N.py in the session workspace; this module bundles
those runs plus the conversation into a self-contained, re-executable .ipynb
with a provenance header (parameter ids, time, library versions).
The saved runs are rewritten to standalone code (to_standalone): load_data()
becomes a fetch_series(...) call that wraps spz.get_data and blanks the declared
fill value exactly as the session did, the agent-only param_card()/
document_method() calls are dropped, and a minimal header supplies the imports
plus real clean()/export()/magnitude()/interp_to() helpers — so each cell
runs in a plain Jupyter kernel with no HelioAI sandbox around it.
export_session_notebook ¶
Write the session as a .ipynb and return its path.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
user_id
|
str
|
Storage owner. |
required |
session_id
|
str
|
Session to export. |
required |
out_path
|
Path | None
|
Destination. Defaults to |
None
|
Returns:
| Type | Description |
|---|---|
Path
|
The path written. |
Example
export_session_notebook("cli", "8f3aa012-...") PosixPath('.../data/users/cli/workspace/plot-imf-bz_8f3aa0/plot-imf-bz_8f3aa0.ipynb')
The notebook opens with a provenance header (parameter ids, library versions), then one runnable cell per saved sandbox run, rewritten to standalone speasy calls.
Source code in helioai/export.py
to_standalone ¶
Turn a saved sandbox run into standalone, re-executable code.
Strips agent-only calls, rewrites load_data() → fetch_series()/fetch_events(), and (unless embedded in a notebook that already has a setup cell) prepends the imports plus the real helpers it needs.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
code_src
|
str
|
A |
required |
manifest
|
dict
|
The session manifest from |
required |
with_header
|
bool
|
Prepend the standalone imports/helpers header. |
True
|
Example
manifest = {"datasets": {"imf_gsm": {"kind": "timeseries", ... "param_id": "amda/imf_gsm", "start": "2005-01-17", "stop": "2005-01-18"}}} to_standalone('d = load_data("imf_gsm")\nparam_card(d, "amda/imf_gsm")\n', ... manifest, with_header=False) 'd = fetch_series("amda/imf_gsm", "2005-01-17", "2005-01-18")'
Source code in helioai/export.py
Configuration¶
helioai.config ¶
Centralized configuration — loads .env once at startup.
settings is a module-level singleton imported everywhere.
Importing this module never requires an API key. Credentials are validated where they
are used, by llm.factory.build_llm_client, which already raised the same errors — the
import-time copy only meant that surfaces needing no LLM at all could not start. The MCP
server is exactly that: a pure tool provider whose client brings its own model, and it
died on AZURE_OPENAI_API_KEY is not set before serving a single tool.
EXPERIMENTS
module-attribute
¶
EXPERIMENTS: frozenset[str] = frozenset({'deferred_tools', 'search_budget', 'search_variables', 'judgment_intent'})
The behaviours that change what the model sees or is told, each off by default.
Every one of them was committed on the strength of a single live run and never measured
against the loop it replaced. They stay in the code as named experiments so that each can
be switched on alone and compared on the same questions, N runs each
(scripts/bench_live.py):
deferred_tools: the lead sees the formulary and catalogue tools only after asking for them withsearch_tools, and its prompt says so.search_budget: adata_analystpast three lookups (aplasma_physicistpast two) with nothing downloaded receives a correction listing the ids it already has.search_variables: the top hit of a parameter search lists every variable of its dataset.judgment_intent: with a judging backend (HELIOAI_JUDGMENT_BACKEND=jev), the user's question is read once by the judge, concurrently with the first model call, into an intent contract — deliverable, quantity, date, whether an uncertainty, a method or two spacecraft were asked for — placed against what the turn loaded and delivered (core/joins.py: frame, window, quantity, responsiveness) and emitted as anintentevent after the answer. Observation: nothing in the loop acts on it.
final_answer was one of them and is now the default: on the third bench (34 clean runs,
four configurations) it cost nothing on any question, re-enabled the claim verdict — zero
contradictions over nine — and named rejected candidates as often as the loop it replaced.
The two search experiments showed no value there (eight id-resolution runs correct without
them); they stay off until a question that fails without them is recorded.
JUDGMENT_BACKENDS
module-attribute
¶
Who answers the judgment questions of helioai.core.judgment.
null abstains on every question, and it is the default: every call site runs the
deterministic code it runs today when the answer is None, so a build with this backend
is the loop as it was before the module existed. jev is TypeSafe's System One model
through the optional judgment extra. There is no third "local" backend on purpose — the
deterministic path is not a judge, it is the fallback, and naming it as a backend would
suggest the two could disagree.
AzureOpenAIConfig
dataclass
¶
Azure OpenAI deployment settings.
Azure routes by deployment name rather than model name, and reasoning models
(GPT-5, o-series) reject an explicit temperature — hence temperature=None
by default, which omits the field entirely.
Source code in helioai/config.py
GeminiConfig
dataclass
¶
Google Gemini settings, used with the native google-genai client.
Source code in helioai/config.py
GroqConfig
dataclass
¶
Groq settings. Reached through the shared OpenAI-compatible client.
Source code in helioai/config.py
OpenCodeConfig
dataclass
¶
OpenCode's Zen gateway — OpenAI-compatible, whichever way you reach it: the flat-rate Go subscription, a BYOK-routed key, or any other model Zen hosts.
base_url defaults to the Go-plan endpoint, since that flat-rate tier is what
most accounts actually have. It is a DIFFERENT catalogue from the general Zen
endpoint (.../zen/v1, no /go/) — that one serves premium/BYOK-only models
(Claude, ...) a Go subscription cannot reach, confirmed by querying both
/models endpoints directly. Override HELIOAI_OPENCODE_URL to the plain Zen
path if your access is not the Go plan.
No default model: what is reachable depends on your plan/BYOK setup and Zen's
rotating catalogue (GLM, Kimi, DeepSeek, Qwen, MiniMax...). Set
HELIOAI_OPENCODE_MODEL to the exact id from your dashboard — an empty string
fails at the API with a clear "unknown model" rather than silently routing to a
guessed default that may not exist on your plan.
16384 output tokens because everything this gateway serves is a reasoning model, and reasoning, prose AND tool-call arguments all draw on the same budget. At 4096, DeepSeek v4's run_python calls (~12k chars of JSON once a plot script is in them) were cut mid-string: the model saw "missing 1 required positional argument: 'code'", could not know why, and burned five turns re-sending the same truncated call. Same lesson as Azure's 2048→8192, one provider later.
Source code in helioai/config.py
OllamaConfig
dataclass
¶
Local Ollama settings.
Ollama serves an OpenAI-compatible API on /v1, so it needs no client of its
own and no API key. Point base_url elsewhere for any other local endpoint.
Source code in helioai/config.py
LLMConfig
dataclass
¶
Which provider to use, and the settings for each one.
opencode is the default because it is what the README tells a new user to set up:
one key reaches hosted reasoning models. The default was
azure until 0.4.0, an enterprise deployment no newcomer has, so the first error
anyone saw named a key they had never heard of.
Source code in helioai/config.py
AgentConfig
dataclass
¶
Agent loop limits, and which model each delegated role runs on.
max_iterations caps how many tool-calling rounds one question may take
before the loop gives up, bounding both runtime and token spend.
role_models maps a sub-agent role to (provider, model). A parameter_hunter
resolves ids from search results and needs no frontier model; a data_analyst
writes the physics and does. Left empty, every role runs on the lead's client, as
it always did. Parsed from HELIOAI_ROLE_MODELS="parameter_hunter=groq:llama-3.3-70b-
versatile,data_analyst=opencode" — the model part is optional and defaults to the
provider's configured model.
experiments names the behaviours of EXPERIMENTS that are switched on, from
HELIOAI_EXPERIMENTS. Empty — the default — is the loop as it behaved before any of
them existed.
Source code in helioai/config.py
RAGConfig
dataclass
¶
Parameter search settings.
Retrieval is hybrid: dense embeddings for descriptions, BM25 for exact tokens
like BGSEc, fused by Reciprocal Rank Fusion with parameter rrf_k.
There is no cross-encoder reranking stage, and that is a measured decision, not
an omission: a generic MS MARCO cross-encoder (ms-marco-MiniLM-L-6-v2) was tried
over the fused candidates and degraded results — trained on web prose, it
discards the dense+sparse consensus that makes exact-code matching work. The
plumbing sat disabled for a year and was removed; only a domain-tuned reranker
would be worth adding back.
hybrid_fetch_k stays at 50 for the same kind of reason. Raising it to 100 or 200,
or the BM25 cap alone to 100, was measured on 2026-09-22 against the 30 HelioBench
n1 queries: MRR fell from 0.736 to 0.724, 0.712 and 0.729, because a wider pool lets
an unpenalised candidate from deep in both lists climb over a relevant one that
carries a small penalty, and two accepted products dropped out of the top-k
altogether. The product the raise was meant to rescue (ssc/mms1, rank 54 in BM25)
is reached through the dense channel instead, once ef_search looks wide enough.
index_repo is the Hugging Face dataset helioai index fetches a prebuilt index
from when the local one is empty (helioai.index_snapshot). It is a setting so that
a lab can point at its own mirror, and an empty value turns fetching off.
Source code in helioai/config.py
WorkspaceConfig
dataclass
¶
Retention of per-session working directories.
Their location is not a setting: workspace.user_home derives it from data_dir
per user. A workspace_dir field (and HELIOAI_WORKSPACE) used to sit here, read
from the environment and consumed by nothing since storage became per-user.
Source code in helioai/config.py
ProfileConfig
dataclass
¶
Legacy location of the single-user profile, kept for helioai migrate-storage.
The agent reads users/<user>/profile.md (workspace.user_home) since storage
became per user; helioai profile, %helioai_profile and the web UI all edit that
file. Nothing injects this path any more, so HELIOAI_PROFILE — which only moved
it — was a knob that did nothing, and is gone. The path stays as the place the
migration looks for a profile written by an older install.
Source code in helioai/config.py
RecipesConfig
dataclass
¶
Where scientific recipes are loaded from.
Defaults to the copy shipped inside the package so pip install works;
override with HELIOAI_RECIPES_DIR to use your own set.
Source code in helioai/config.py
CatalogsConfig
dataclass
¶
LiteratureConfig
dataclass
¶
MCPConfig
dataclass
¶
Remote MCP servers to mount, plus auth for HelioAI's own MCP HTTP transport.
Source code in helioai/config.py
VisionConfig
dataclass
¶
Multimodal review of generated figures.
A stateless side-call outside the agent loop: the image is downscaled, sent once, and only the text verdict enters the history — never the image, which would otherwise be resent on every subsequent turn. Off by default.
Source code in helioai/config.py
JudgmentConfig
dataclass
¶
A System One judge beside the loop, in observation.
Every question HelioAI asks it lives in helioai.core.judgment, so what is delegated
to a proprietary model can be audited in one file. The backend is one axis; which
sites may ask is the other, one EXPERIMENTS name per site — so "the judge does not
help here" and "the layer costs something" can be told apart. Off by default, and
nothing it answers corrects the model: it annotates and it is recorded.
Source code in helioai/config.py
DevConfig
dataclass
¶
Shared secret unlocking unrestricted mode past the heliophysics guardrail.
Empty by default, which means no token is valid and every request stays scoped. Compared in constant time.
Source code in helioai/config.py
WebAuthConfig
dataclass
¶
Nominative tokens for the web UI, parsed from HELIOAI_USERS.
Empty means no authentication and a single local user, which is the intended behaviour for local development only.
Source code in helioai/config.py
Settings
dataclass
¶
Root settings object.
Imported as the module-level settings singleton and read everywhere; built
once at import by _load(). Credentials are not checked here — importing must
work with no key at all — but in llm.factory.build_llm_client, where the
selected provider is actually used.
Source code in helioai/config.py
validate_experiments ¶
Refuse unknown experiment names — an error, not a warning.
An experiment that silently does nothing would be measured as if it did, and the comparison would be wrong without anyone knowing.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
experiments
|
frozenset[str] | None
|
The set to check; |
None
|
Returns:
| Type | Description |
|---|---|
frozenset[str]
|
The same set, when every name is known. |
Raises:
| Type | Description |
|---|---|
RuntimeError
|
Naming the unknown entries and the known ones. |
Source code in helioai/config.py
validate_judgment ¶
Refuse an unknown judgment backend — an error, not a warning.
Same reasoning as validate_experiments: a backend name that silently fell back to
abstaining would be measured as if it had judged. Called where the API key is
already checked, never at import.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
cfg
|
JudgmentConfig | None
|
The config to check; |
None
|
Returns:
| Type | Description |
|---|---|
str
|
The backend name, when known. |
Raises:
| Type | Description |
|---|---|
RuntimeError
|
Naming the unknown backend and the known ones. |
Source code in helioai/config.py
dev_unlock ¶
True iff the supplied token matches the configured dev secret.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
supplied
|
str | None
|
Token offered by the caller — a |
required |
Returns:
| Type | Description |
|---|---|
bool
|
True only when a dev token is configured and the supplied one matches. |
bool
|
An unconfigured instance answers False for every input, including |
bool
|
|
Source code in helioai/config.py
Logging¶
helioai.logging_config ¶
Structured logging via structlog.
Output format is selected by HELIOAI_LOG_FORMAT
console(default): human-friendly, colourised.json: one JSON object per line.
setup_logging ¶
Configure structlog and the root logger.
Output format follows HELIOAI_LOG_FORMAT: console (default) or json.
Safe to call more than once — every entry point calls it, and repeated calls
replace the handler rather than stacking duplicates.
HELIOAI_LOG_LEVEL overrides level. Every entry point hardcodes its own,
so without this there is no way to quiet a third party that logs at the same
level — speasy's inventory probes warn loudly on a provider it then disables,
which is noise in a recorded session or a demo. An unrecognised value is
ignored rather than obeyed: a typo must not silently turn logging up.
tracebacks=False is for the terminal a person is reading: an error keeps its one
log line, with the exception's type and message, but not the stack under it — the
interface prints what to do instead. DEBUG brings the stack back, whatever the flag.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
level
|
str | int
|
Log level name or numeric value. Unknown names fall back to INFO. |
'INFO'
|
tracebacks
|
bool
|
Render the stack of a logged exception. |
True
|
Source code in helioai/logging_config.py
get_logger ¶
Return a structlog logger, optionally bound to a module name.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
name
|
str | None
|
Usually |
None
|
Returns:
| Type | Description |
|---|---|
Any
|
A structlog bound logger. Its output format follows |
Any
|
|
Any
|
rather than here. |