Skip to content

Add current local-runtime and inference wire contracts - #3

Merged
senamakel merged 1 commit into
mainfrom
runtime-metadata-and-latest
Aug 30, 2026
Merged

Add current local-runtime and inference wire contracts#3
senamakel merged 1 commit into
mainfrom
runtime-metadata-and-latest

Conversation

@senamakel

Copy link
Copy Markdown
Member

Summary

  • add runtime-neutral reasoning, cache, continuation, and provider metadata needed by current TinyAgents
  • preserve current local runtime probing, warm-up, model validation, and Ollama embedding behavior
  • fix strict schema fallback, tool id normalization, cache accounting, Responses request/reasoning handling, and JSON Schema union validation
  • keep tracing test-only and use httpdate instead of chrono

Verification

  • cargo test
  • cargo clippy --all-targets -- -D warnings
  • downstream: cargo test --workspace
  • downstream: cargo clippy --workspace --all-targets -- -D warnings

Co-authored-by: Medulla <medulla@tinyhumans.ai>
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Aug 30, 2026

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review Completed 2026-08-30T17:59:12.023682Z 70412ce PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@coderabbitai

coderabbitai Bot commented Aug 30, 2026

Copy link
Copy Markdown

Warning

Review limit reached

  • Run on-demand review

On-demand reviews are free for the next 21 days. After that, they cost $0.25 per reviewed file.

Or wait 8 minutes for your next included review.

View limit details

Limit details: You’ve used the included review currently available.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 4c6cd872-9fcd-4f14-9b4d-cbf8d31df189

📥 Commits

Reviewing files that changed from the base of the PR and between 2c40a5f and 70412ce.

📒 Files selected for processing (20)
  • crates/tinyinference/src/cache/mod.rs
  • crates/tinyinference/src/cache/types.rs
  • crates/tinyinference/src/embeddings/mod.rs
  • crates/tinyinference/src/embeddings/ollama.rs
  • crates/tinyinference/src/message/mod.rs
  • crates/tinyinference/src/message/types.rs
  • crates/tinyinference/src/model/mod.rs
  • crates/tinyinference/src/model/types.rs
  • crates/tinyinference/src/providers/mock.rs
  • crates/tinyinference/src/providers/openai/convert.rs
  • crates/tinyinference/src/providers/openai/local.rs
  • crates/tinyinference/src/providers/openai/local_test.rs
  • crates/tinyinference/src/providers/openai/mod.rs
  • crates/tinyinference/src/providers/openai/responses.rs
  • crates/tinyinference/src/providers/openai/sse.rs
  • crates/tinyinference/src/providers/openai/test.rs
  • crates/tinyinference/src/providers/openai/transport.rs
  • crates/tinyinference/src/providers/openai/types.rs
  • crates/tinyinference/src/providers/types.rs
  • crates/tinyinference/src/tool.rs

Warning

Your free Security trial is over. An organization admin can activate Security or dismiss this notice.


Comment @coderabbitai help to get the list of available commands.

@tinysweeper

tinysweeper Bot commented Aug 30, 2026

Copy link
Copy Markdown

How this change flows

0 changed behaviours across 2 relationships. 2 surrounding behaviours are shown (60 graph nodes walked). 49 further behaviours left out to keep the diagram readable.

flowchart LR
  n0["translate_request"]:::impacted
  n1["translate_request_with"]:::impacted
  n0 -->|calls| n1
  n0 -->|tests| n1
  classDef changed fill:#0d4429,stroke:#238636,color:#e6edf3
  classDef impacted fill:#161b22,stroke:#6e7681,color:#c9d1d9
  classDef flagged fill:#5a1e02,stroke:#d93f0b,color:#ffffff
  classDef blocking fill:#67060c,stroke:#f85149,color:#ffffff
Loading

Green: changed behaviour. Grey: surrounding behaviour. Arrows name the call, use, implementation, or test relationship. Orange: has findings. Red: has a finding that blocks the merge.

tinysweeper 0.1.0

@tinysweeper tinysweeper Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tinysweeper found nothing blocking. Approving.

$0.0000 · 0 in / 0 out

@senamakel
senamakel merged commit cc8aca4 into main Aug 30, 2026
13 checks passed

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 70412ce592

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +1415 to +1416
if degrade.native_tools {
self.native_tools_on_wire.store(false, Ordering::Relaxed);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Parse prompt-guided calls after degrading native tools

When a model profile advertises tool calling but the server returns a 400 saying tools are unsupported, this latch makes the retry use prompt-guided <tool_call> output. The response paths at invoke and stream, however, only call prompt_tools::apply_to_response when self.profile.tool_calling is false; the profile remains true here, so the retry can succeed while returning the tool-call markup as ordinary assistant text instead of a normalized call. Base response parsing on this latch as well as the static profile.

AGENTS.md reference: AGENTS.md:L43-L45

Useful? React with 👍 / 👎.

} else {
Vec::new()
};
let tool_choice = (!tools.is_empty()).then(|| translate_tool_choice(&request.tool_choice));

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Use the Responses shape for named tool choice

For ToolChoice::Tool, this reuses the Chat Completions translator and serializes {"type":"function","function":{"name":...}}. The Responses API uses the flattened {"type":"function","name":...} shape, which the removed responses_tool_choice helper previously produced, so named-tool requests on the Responses path are rejected instead of forcing the requested tool.

Useful? React with 👍 / 👎.

Comment on lines +470 to 475
ModelResponse {
message: AssistantMessage {
id: None,
content: vec![ContentBlock::Text(text)],
tool_calls,
content,
tool_calls: Vec::new(),
usage,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Decode Responses function calls

Whenever /responses returns a function_call output item, this now unconditionally emits an empty tool_calls list; the response wire struct also no longer retains the call id, name, or arguments. The previous parser normalized these items, including malformed arguments, so tool-using Responses calls now appear to consumers as empty assistant replies and the requested tool is never executed.

AGENTS.md reference: AGENTS.md:L43-L45

Useful? React with 👍 / 👎.

Comment on lines +312 to +313
Message::User(m) => message_text(&m.content),
Message::Assistant(m) => message_text(&m.content),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Preserve images in Responses input

For a Responses-primary request containing ContentBlock::Image, message_text retains only text blocks, so this branch silently discards every image; an image-only user turn is skipped altogether at the subsequent empty-text check. The prior translation emitted input_image content parts, so vision requests now run against missing input and can return plausible but incorrect results rather than an error.

Useful? React with 👍 / 👎.

} else {
"stop".to_string()
}),
finish_reason: Some("stop".to_string()),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Preserve incomplete Responses finish reasons

Every Responses result is now marked stop, including responses whose status is incomplete because max_output_tokens was reached. Consumers therefore cannot distinguish truncated output from a completed answer and may accept or cache partial text/JSON; retain the response status and map incomplete_details.reason as the previous parser did.

AGENTS.md reference: AGENTS.md:L43-L45

Useful? React with 👍 / 👎.

Comment on lines +186 to +191
model_info
.iter()
.filter(|(key, _)| key.ends_with(".context_length") || key.as_str() == "context_length")
.filter_map(|(_, value)| value.as_u64())
.filter(|value| *value > 0)
.min()

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Select the language-model context length

On multimodal Ollama models, model_info can contain both the language architecture window (for example gemma3.context_length) and a small projector window such as clip.context_length = 77. Taking the minimum advertises 77 tokens as the chat model's input capacity, causing capability checks or context compaction to reject or discard nearly every normal prompt; identify the main architecture's context field instead of minimizing unrelated components.

Useful? React with 👍 / 👎.

Comment on lines 343 to +344
if live.is_empty() {
return Err(Error::Validation(
"Ollama embedding batches must not contain blank inputs".into(),
));
return Ok(vec![Vec::new(); texts.len()]);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Reject blank embedding batches instead of returning zero-width vectors

For any nonzero-dimensional OllamaEmbeddingModel, an all-blank batch now returns one empty vector per input, violating EmbeddingModel's fixed-dimension contract. Passing such output through Retriever::index immediately fails in InMemoryVectorStore::add, while direct callers can accidentally persist invalid vectors elsewhere; return a validation error or valid vectors of dimensions() instead.

AGENTS.md reference: AGENTS.md:L53-L58

Useful? React with 👍 / 👎.

Comment on lines +452 to +457
let parsed: ResponsesResponse =
serde_json::from_value(value.clone()).unwrap_or_else(|_| ResponsesResponse {
output: Vec::new(),
output_text: None,
usage: None,
});

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Propagate invalid Responses payload shapes

If a successful HTTP response has an incompatible schema, such as {"output":"not-an-array"}, deserialization now falls back to an empty response and invoke_responses returns a successful blank assistant message. This hides provider incompatibilities and malformed payloads that previously surfaced as serialization errors, making failures indistinguishable from genuine empty completions; keep parsing fallible and propagate the decode error.

Useful? React with 👍 / 👎.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant