Files
homelab-frontend/LLM_TOOL_CALLS.md
T
Story Crater Bot b0ce2fb67c feat: build and publish the gateway image via Forgejo Actions
- Dockerfile: multi-stage, distroless nonroot, CGO_ENABLED=0 static, commit
  SHA stamped via VERSION build arg.
- .forgejo/workflows/ci.yaml: Forgejo reads .forgejo/, not .github/, and the
  runner declares only the "docker" label. Verify job on every push; image
  build and push gated to main.
- Drop .github/workflows/ci.yml — this remote is Forgejo, so it never ran.
- deployment.yaml: image from the Forgejo registry, forgejo-registry pull
  secret, runAsUser 65532 to match distroless nonroot.
- kustomization.yaml: pin the tag in one place. Promoting a build is a
  one-line newTag bump, never :latest.
2026-08-19 21:48:11 -07:00

12 KiB

LLM Tool Calls Testing Guide

This guide shows how to test the gateway with LLM tool calling (function calling) across different APIs.

What's Tested

The gateway fully supports tool calling for:

  • OpenAI API (/v1/chat/completions) - OpenAI, DeepSeek, etc.
  • Anthropic API (/llm/v1/messages) - Claude models
  • Custom APIs - Any LLM that supports tool definitions and responses

Test Coverage

✅ OpenAI-style tool calling
✅ Streaming tool calls (SSE with tool_use blocks)
✅ Multi-turn conversations with tool results
✅ Parallel tool calls (multiple tools at once)
✅ Anthropic tool_use format
✅ Complex nested tool arguments

Quick Start: Run Tests Locally

cd /Users/rockliang/workplace/homelab-frontend

# Run all tool call tests
go test ./internal/proxy/... -run "Tool" -v

# Or run with race detector (recommended)
go test -race ./internal/proxy/... -run "Tool" -v

# Expected output: 6 tests, all passing

Test Scenarios

1. OpenAI-Style Tool Calling

What it tests:

  • Request with tool definitions reaches upstream unmodified
  • Upstream can return tool_calls in response
  • Response with tool_calls passes through to client

Test code:

// Request
{
  "model": "reasoning",
  "messages": [{"role": "user", "content": "What's the weather?"}],
  "tools": [{
    "type": "function",
    "function": {
      "name": "get_weather",
      "description": "Get weather for a location",
      "parameters": {...}
    }
  }]
}

// Response (from upstream)
{
  "choices": [{
    "message": {
      "tool_calls": [{
        "id": "call_abc123",
        "type": "function",
        "function": {
          "name": "get_weather",
          "arguments": "{\"location\":\"San Francisco\"}"
        }
      }]
    }
  }]
}

Run:

go test ./internal/proxy/... -run TestToolCallOpenAIStyle -v

2. Streaming Tool Calls

What it tests:

  • Tool calls can be streamed (SSE format)
  • Multiple chunks arrive with tool_call deltas
  • Stream completes with [DONE] sentinel

Test code:

Chunk 1: {"delta": {"role": "assistant"}, ...}
Chunk 2: {"delta": {"tool_calls": [{"id": "call_123", "function": {...}}]}, ...}
Chunk 3: {"delta": {}, "finish_reason": "tool_calls"}
Chunk 4: [DONE]

Run:

go test ./internal/proxy/... -run TestToolCallStreaming -v

3. Multi-Turn Conversation with Tool Results

What it tests:

  • Client can send previous assistant's tool_calls back
  • Tool result can be sent as a "tool" role message
  • Assistant responds with final answer using tool result

Flow:

Turn 1: User asks → LLM decides to call tool
Turn 2: Client sends tool result → LLM generates final answer

Test code:

// Turn 1 Request
{
  "model": "reasoning",
  "messages": [{"role": "user", "content": "What's the weather?"}],
  "tools": [...]
}

// Turn 1 Response (tool_calls)
{
  "message": {
    "tool_calls": [{
      "id": "call_abc",
      "function": {"name": "get_weather", "arguments": "..."}
    }]
  }
}

// Turn 2 Request (with tool result)
{
  "model": "reasoning",
  "messages": [
    {"role": "user", "content": "What's the weather?"},
    {"role": "assistant", "tool_calls": [...]},
    {"role": "tool", "content": "{\"temperature\": 22, \"condition\": \"cloudy\"}"}
  ],
  "tools": [...]
}

// Turn 2 Response (final answer)
{
  "message": {
    "content": "The weather in San Francisco is 22°C and cloudy."
  }
}

Run:

go test ./internal/proxy/... -run TestToolCallMultiTurn -v

4. Parallel Tool Calls

What it tests:

  • LLM can request multiple tools in one response
  • Gateway preserves all tool_calls
  • Client can execute them in parallel

Test code:

// Single response with 3 tool_calls
{
  "message": {
    "tool_calls": [
      {"id": "call_1", "function": {"name": "get_weather", "arguments": "{\"location\":\"New York\"}"}},
      {"id": "call_2", "function": {"name": "get_weather", "arguments": "{\"location\":\"London\"}"}},
      {"id": "call_3", "function": {"name": "get_weather", "arguments": "{\"location\":\"Tokyo\"}"}}
    ]
  }
}

Run:

go test ./internal/proxy/... -run TestParallelToolCalls -v

5. Anthropic Tool Use Format

What it tests:

  • Different tool format: Anthropic uses tool_use blocks instead of tool_calls
  • Gateway handles both formats transparently
  • Tools are sent with tools parameter

OpenAI format:

{"tool_calls": [{"type": "function", "function": {...}}]}

Anthropic format:

{"content": [
  {"type": "text", "text": "..."},
  {"type": "tool_use", "id": "...", "name": "...", "input": {...}}
]}

Test code:

// Request
{
  "model": "claude",
  "messages": [{"role": "user", "content": "What's the weather?"}],
  "tools": [{
    "name": "get_weather",
    "description": "Get weather",
    "input_schema": {...}
  }]
}

// Response (Anthropic format)
{
  "content": [
    {"type": "text", "text": "I'll check the weather..."},
    {"type": "tool_use", "id": "toolu_123", "name": "get_weather", "input": {...}}
  ],
  "stop_reason": "tool_use"
}

Run:

go test ./internal/proxy/... -run TestAnthropicToolUse -v

6. Complex Nested Tool Arguments

What it tests:

  • Tool arguments can be complex JSON structures
  • Nested objects, arrays, and deeply nested data preserved
  • No argument modification or parsing

Test code:

{
  "function": {
    "name": "create_event",
    "arguments": {
      "title": "Team Meeting",
      "time": "2025-08-20T14:00:00Z",
      "attendees": [
        {"name": "Alice", "email": "[email protected]"},
        {"name": "Bob", "email": "[email protected]"}
      ],
      "location": {
        "address": "123 Main St",
        "city": "San Francisco",
        "country": "USA"
      },
      "tags": ["important", "recurring"]
    }
  }
}

Run:

go test ./internal/proxy/... -run TestComplexToolArguments -v

Running Against Real LLMs

With Local Stubs (Current)

Tests use mock HTTP servers, so they run instantly:

go test ./internal/proxy/... -run "Tool" -v
# All 6 tests complete in ~220ms

With Real Upstreams (Future)

Once you have real LLM services running, update the config:

# k8s/configmap.yaml
models:
  - name: "reasoning"
    address: "reasoning-predictor.llm-serving:80"  # Real upstream
  - name: "claude"
    address: "claude-api.anthropic.com:443"         # Real Anthropic

Then use the gateway normally:

# Terminal 1: Start gateway
export CONFIG_PATH=config.yaml
go run ./cmd/gateway

# Terminal 2: Test with real LLM
curl -X POST http://localhost:8080/v1/chat/completions \
  -H 'content-type: application/json' \
  -d '{
    "model": "reasoning",
    "messages": [{"role": "user", "content": "What color is the sky?"}],
    "tools": [{
      "type": "function",
      "function": {
        "name": "search",
        "description": "Search the internet",
        "parameters": {"type": "object", "properties": {}}
      }
    }]
  }'

Gateway Behavior with Tool Calls

Request Path

Client Request
    ↓
Body-based dispatch (find model)
    ↓
Look up upstream address
    ↓
Forward request unmodified (including tools)
    ↓
Upstream LLM processes tools

Response Path

Upstream Response (with tool_calls or tool_use)
    ↓
Stream unbuffered if streaming
    ↓
Return to client exactly as received

Key Properties

  1. No Rewriting: Tool definitions and responses pass through unmodified
  2. Format Agnostic: Both OpenAI tool_calls and Anthropic tool_use work
  3. Streaming Safe: Tool calls stream incrementally without buffering
  4. Nested Structures: Complex JSON arguments fully preserved

Common Tool Call Patterns

Pattern 1: Sequential Tool Use

Client → LLM (please use search tool)
       ← LLM (tool_calls: [search(...)])
Client → (execute search, send results)
Client → LLM (here are search results)
       ← LLM (final answer)

Pattern 2: Parallel Tool Calls

Client → LLM (check weather in 3 cities)
       ← LLM (tool_calls: [get_weather(NY), get_weather(LA), get_weather(SF)])
Client → (execute all 3 in parallel)
Client → LLM (here are all results)
       ← LLM (summary)

Pattern 3: Tool Result Formatting

Client receives tool_calls with:
- id: unique identifier
- function.name: tool name
- function.arguments: JSON string (always a string, not parsed object)

Client sends back:
- role: "tool"
- content: result JSON string
- tool_call_id: matches the original call id

Verification Checklist

  • OpenAI-style tool definitions forward to upstream
  • Tool calls in response reach client unmodified
  • Streaming tool calls arrive incrementally
  • Multi-turn conversations preserve tool context
  • Parallel tool calls all included in response
  • Anthropic tool_use format works
  • Complex nested arguments preserved

Run all:

go test ./internal/proxy/... -run "Tool" -v --race

Expected: 6/6 passing, race detector clean


Integration with Other Phases

Phase 2.9: Anthropic Dialect

Currently, Anthropic tool calls work through the generic route handler. Phase 2.9 will add a dedicated /llm/v1/messages endpoint with full Anthropic-specific handling.

Phase 2.13: Error Handling

Tool call errors (unknown tool, parsing errors) will have proper error responses in both OpenAI and Anthropic formats.

Phase 3: Authentication

Tool calls work with all authentication methods (bearer tokens, API keys) - no special handling needed since tools are just part of the message payload.

Phase 4: Rate Limiting

Tool calling counts the same as regular chat requests. Rate limits apply per conversation, not per tool call.


Debugging Tool Calls

Check if tool definitions reach upstream:

# Enable request logging
go test ./internal/proxy/... -run TestToolCallOpenAIStyle -v 2>&1 | grep -A5 "tool"

Verify tool response format:

# Extract and pretty-print response
curl -X POST http://localhost:8080/v1/chat/completions ... | jq '.choices[0].message.tool_calls'

Test streaming tool calls:

curl -N http://localhost:8080/v1/chat/completions \
  -H 'content-type: application/json' \
  -d '{..., "stream": true, "tools": [...]}'
# Should see incremental chunks with tool_use deltas

FAQ

Q: Do I need to modify the gateway code to support tool calls? A: No. Tool calls are just JSON in the request/response body. The gateway forwards them unchanged.

Q: What if the LLM doesn't support tools? A: The tool definitions are simply ignored. The gateway doesn't validate or enforce tool support.

Q: Can I mix OpenAI and Anthropic tool formats? A: Not in the same request. OpenAI clients expect tool_calls, Anthropic clients expect tool_use blocks. The upstream API determines the format.

Q: How are tool arguments limited? A: By the per-route maxBodySize config. Complex nested arguments count toward that limit.

Q: Can tool calls be streamed? A: Yes! SSE streaming fully supports tool calls. They arrive in delta chunks like text tokens.


Next Steps

  1. Run tests locally:

    go test ./internal/proxy/... -run "Tool" -v
    
  2. Deploy to cluster: See CLUSTER_REPO_SETUP.md

  3. Test against real LLMs: Update config with real upstream addresses, restart gateway

  4. Phase 2.9: Implement Anthropic dialect handler for /llm/v1/messages

  5. Phase 4: Add tool call budgeting and rate limits