747 lines
23 KiB
Markdown
747 lines
23 KiB
Markdown
# Poimen Workflows
|
|||
|
|
|
||
|
|
**Temporal-powered orchestration that transforms natural language into durable, scalable workflow executions.**
|
||
|
|
|
||
|
|
[](https://github.com/your-org/workflows/actions/workflows/ci.yml)
|
||
|
|
[](https://goreportcard.com/report/github.com/your-org/workflows)
|
||
|
|
[](./docs/coverage.md)
|
||
|
|
|
||
|
|
## What is Poimen?
|
||
|
|
|
||
|
|
**Poimen** (Greek: Ποιμήν) means **"shepherd"** — a guide who tends, orchestrates, and reconciles. In this project, Poimen acts as an intelligent **orchestration shepherd** for distributed agent workflows:
|
||
|
|
|
||
|
|
- 🐑 **Shepherds workflows** — Guides activities from intent to completion
|
||
|
|
- 🔄 **Reconciles state** — Ensures consistency across distributed steps
|
||
|
|
- 🎯 **Routes intelligently** — Uses LLM reasoning to choose the best path forward
|
||
|
|
- 💾 **Remembers lessons** — Semantic memory prevents repeated mistakes
|
||
|
|
- 🛡️ **Fault-tolerant** — Temporal's durable execution keeps workflows safe from failure
|
||
|
|
|
||
|
|
Like a shepherd managing a flock, Poimen manages orchestrated agent deployments at scale—keeping tasks aligned, recovering from chaos, and learning from each execution.
|
||
|
|
|
||
|
|
## Overview
|
||
|
|
|
||
|
|
Poimen Workflows orchestrates complex agent-driven tasks using **LLM-powered routing**, **semantic memory integration**, and **durable state machines**. Write your intent in plain English—the system generates executable Temporal workflows, manages retries and error handling, and scales to thousands of concurrent executions.
|
||
|
|
|
||
|
|
**Core Promise:** Any activity (code scan, deployment, notification) becomes a composable building block in a larger reconciliation pipeline, with intelligent ordering, memory-augmented context, and automatic fault recovery.
|
||
|
|
|
||
|
|
---
|
||
|
|
|
||
|
|
## Architecture & Data Flow
|
||
|
|
|
||
|
|
```
|
||
|
|
User Request
|
||
|
|
↓
|
||
|
|
├─→ [1. LLM Router]
|
||
|
|
│ • Analyze intent & extract parameters
|
||
|
|
│ • Query semantic memory for domain knowledge
|
||
|
|
│ • Build WorkflowSpec with activity graph
|
||
|
|
│ • Inject retry policies & error handlers
|
||
|
|
│
|
||
|
|
├─→ [2. Memory Retrieval]
|
||
|
|
│ • RetrieveMemoryActivity queries poimen-memory (Rust semantic search)
|
||
|
|
│ • Returns relevant skills, lessons, best practices
|
||
|
|
│ • Augments LLM prompts with domain context
|
||
|
|
│
|
||
|
|
├─→ [3. Workflow Validation]
|
||
|
|
│ • StateGraphValidator: Check DAG structure
|
||
|
|
│ • ActivityAvailabilityValidator: Verify all activities registered
|
||
|
|
│ • TimeoutValidator: Validate timeout formats
|
||
|
|
│ • OutputMatchValidator: Ensure parameter bindings exist
|
||
|
|
│
|
||
|
|
├─→ [4. Temporal Execution]
|
||
|
|
│ • RoutingWorkflow executes state machine
|
||
|
|
│ • JSONPath parameter chaining: ${Step1.output.path}
|
||
|
|
│ • Automatic retries for transient failures
|
||
|
|
│ • Catch blocks for graceful error recovery
|
||
|
|
│ • Durable execution guarantees (fault-tolerant)
|
||
|
|
│
|
||
|
|
└─→ Completed Workflow
|
||
|
|
• Results stored in activity output
|
||
|
|
• Memory updated with execution logs
|
||
|
|
• Notifications sent via DeploymentPreCheckActivity
|
||
|
|
```
|
||
|
|
|
||
|
|
---
|
||
|
|
|
||
|
|
## Philosophy: The Shepherd's Role
|
||
|
|
|
||
|
|
Poimen embodies three core principles:
|
||
|
|
|
||
|
|
1. **Orchestration as Shepherding**
|
||
|
|
- Not rigid task scheduling, but intelligent guidance
|
||
|
|
- Activities work together toward a common goal
|
||
|
|
- The system learns and adapts from each execution
|
||
|
|
|
||
|
|
2. **Reconciliation as Healing**
|
||
|
|
- Workflows reconcile state across distributed systems
|
||
|
|
- Errors are opportunities to learn, not failures to hide
|
||
|
|
- Memory persists lessons so future workflows are smarter
|
||
|
|
|
||
|
|
3. **Scale through Composition**
|
||
|
|
- One shepherd (Poimen) tends infinite flocks (workflows)
|
||
|
|
- Skills compose into larger behaviors
|
||
|
|
- Knowledge compounds—each registration makes routing better
|
||
|
|
|
||
|
|
**The metaphor:** Just as a shepherd doesn't dictate each sheep's movement but guides the flock toward pasture, Poimen doesn't hard-code workflows but learns from context, remembers lessons, and guides agents through complex, adaptive orchestration.
|
||
|
|
|
||
|
|
---
|
||
|
|
|
||
|
|
## Key Features
|
||
|
|
|
||
|
|
### 🧠 LLM-Powered Routing
|
||
|
|
- **Natural Language Input:** `"Scan repo for security issues and notify team"`
|
||
|
|
- **Intelligent Workflow Generation:** Reasoning model generates executable WorkflowSpec
|
||
|
|
- **Activity Knowledge Base:** 9 registered activities inform LLM about timeouts, retry policies, dependencies
|
||
|
|
- **Error Handling Strategies:** Automatic catch blocks and fallback flows
|
||
|
|
|
||
|
|
### 💾 Memory-Augmented Context
|
||
|
|
- **Semantic Search Integration:** Query poimen-memory before routing decisions
|
||
|
|
- **Domain Knowledge Injection:** Skills and lessons automatically augment LLM prompts
|
||
|
|
- **Skill Reusability:** Registered knowledge base prevents redundant work
|
||
|
|
- **Learning Loop:** Execution results feed back into memory for continuous improvement
|
||
|
|
|
||
|
|
### 🔄 Generic State Machine Executor
|
||
|
|
- **JSON WorkflowSpec:** Any activity graph becomes executable
|
||
|
|
- **Parameter Chaining:** JSONPath-based data flow (`${ActivityA.output.field}`)
|
||
|
|
- **Automatic Retries:** Configurable retry policies per activity
|
||
|
|
- **Catch Blocks:** Error recovery with fallback states
|
||
|
|
- **Durable Guarantees:** Temporal ensures exactly-once execution + resume-on-failure
|
||
|
|
|
||
|
|
### 🔌 Extensible Architecture
|
||
|
|
- **LLMProvider Interface:** Swap providers (OpenAI, Claude, local models) at runtime
|
||
|
|
- **SpecBuilder Interface:** Multiple strategies for workflow generation (default, cron, custom)
|
||
|
|
- **WorkflowValidator Interface:** Composable validation pipeline
|
||
|
|
- **ActivityExecutor Interface:** Pluggable execution strategies
|
||
|
|
- **Decorator Pattern:** Caching & retry wrappers for any provider
|
||
|
|
|
||
|
|
---
|
||
|
|
|
||
|
|
## Registered Activities (Knowledge Base)
|
||
|
|
|
||
|
|
| Activity | Purpose | Timeout | Retry |
|
||
|
|
|----------|---------|---------|-------|
|
||
|
|
| **CloneRepo** | Clone git repository | 30s | 3x |
|
||
|
|
| **AnalyzeCode** | Static analysis (SAST) | 120s | 2x |
|
||
|
|
| **SecurityScan** | Dependency & vulnerability scan | 60s | 2x |
|
||
|
|
| **GenerateReport** | Compile results into report | 30s | 1x |
|
||
|
|
| **DeploymentPreCheck** | Validation before deployment | 45s | 2x |
|
||
|
|
| **NotifyStatus** | Send notifications (Slack, email) | 15s | 1x |
|
||
|
|
| **ApproveWorkflow** | Manual approval gate | 3600s | 0x |
|
||
|
|
| **ArchiveResults** | Store results to S3/KV | 30s | 2x |
|
||
|
|
| **RetrieveMemory** | Query semantic memory service | 10s | 1x |
|
||
|
|
|
||
|
|
👉 See [`internal/routing/activity_knowledge_base.json`](./internal/routing/activity_knowledge_base.json) for detailed specs.
|
||
|
|
|
||
|
|
---
|
||
|
|
|
||
|
|
## Getting Started
|
||
|
|
|
||
|
|
### Prerequisites
|
||
|
|
- Go 1.21+
|
||
|
|
- Temporal Server (local or cloud)
|
||
|
|
- Docker (for running Temporal locally)
|
||
|
|
|
||
|
|
### Installation
|
||
|
|
|
||
|
|
```bash
|
||
|
|
git clone https://github.com/your-org/workflows.git
|
||
|
|
cd workflows
|
||
|
|
go mod download
|
||
|
|
```
|
||
|
|
|
||
|
|
### Running Locally
|
||
|
|
|
||
|
|
**1. Start Temporal (via Docker Compose):**
|
||
|
|
```bash
|
||
|
|
docker-compose up -d temporal
|
||
|
|
```
|
||
|
|
|
||
|
|
**2. Register activities and workflow:**
|
||
|
|
```bash
|
||
|
|
cd cmd/worker
|
||
|
|
go build -o worker
|
||
|
|
./worker
|
||
|
|
```
|
||
|
|
|
||
|
|
**3. Trigger a workflow via CLI:**
|
||
|
|
```bash
|
||
|
|
cd cmd/starter
|
||
|
|
go build -o starter
|
||
|
|
./starter --route "Analyze repository for security vulnerabilities"
|
||
|
|
```
|
||
|
|
|
||
|
|
**4. Or trigger programmatically:**
|
||
|
|
```go
|
||
|
|
import "github.com/your-org/workflows/pkg/client"
|
||
|
|
|
||
|
|
ctx := context.Background()
|
||
|
|
c := client.NewClient(...)
|
||
|
|
|
||
|
|
run, err := c.ExecuteWorkflow(ctx, &client.ExecuteWorkflowOptions{
|
||
|
|
ID: "security-scan-1",
|
||
|
|
TaskQueue: "routing",
|
||
|
|
}, "RoutingWorkflow", &routing.WorkflowInput{
|
||
|
|
UserRequest: "Scan repo X for CVEs",
|
||
|
|
})
|
||
|
|
// Wait for result
|
||
|
|
var result routing.WorkflowOutput
|
||
|
|
run.Get(ctx, &result)
|
||
|
|
```
|
||
|
|
|
||
|
|
---
|
||
|
|
|
||
|
|
## Usage Patterns
|
||
|
|
|
||
|
|
### Pattern 1: Simple CLI Routing
|
||
|
|
```bash
|
||
|
|
./starter --route "Deploy service to staging"
|
||
|
|
```
|
||
|
|
|
||
|
|
### Pattern 2: Direct Spec Execution
|
||
|
|
```bash
|
||
|
|
./starter --spec workflow.json
|
||
|
|
```
|
||
|
|
|
||
|
|
### Pattern 3: Service Integration
|
||
|
|
```go
|
||
|
|
// From your application code
|
||
|
|
client := temporal.NewClient()
|
||
|
|
result, err := client.ExecuteWorkflow(ctx, options, "RoutingWorkflow", input)
|
||
|
|
// Implements durable retry + failure recovery
|
||
|
|
```
|
||
|
|
|
||
|
|
### Pattern 4: Custom Validator Chain
|
||
|
|
```go
|
||
|
|
config := &routing.LLMRouterConfig{
|
||
|
|
Validators: []routing.WorkflowValidator{
|
||
|
|
routing.NewStateGraphValidator(),
|
||
|
|
routing.NewActivityAvailabilityValidator(),
|
||
|
|
routing.NewTimeoutValidator(),
|
||
|
|
routing.NewOutputMatchValidator(),
|
||
|
|
},
|
||
|
|
}
|
||
|
|
router := routing.NewLLMRouter(config)
|
||
|
|
```
|
||
|
|
|
||
|
|
---
|
||
|
|
|
||
|
|
## Registerable Skills & Knowledge
|
||
|
|
|
||
|
|
### The Shepherd's Knowledge
|
||
|
|
|
||
|
|
Just as a good shepherd knows the terrain, remembers which paths work, and learns from past journeys, Poimen's intelligence comes from **accumulated knowledge**:
|
||
|
|
|
||
|
|
- **What can be done?** (Activity Knowledge) — Registry of available skills
|
||
|
|
- **How has it been done well?** (Domain Knowledge) — Lessons from past executions
|
||
|
|
- **What patterns work?** (Workflow Patterns) — Proven sequences and error recovery
|
||
|
|
- **What went wrong before?** (Error Knowledge) — Lessons to avoid mistakes
|
||
|
|
- **How should recovery work?** (Execution Policies) — Graceful fallbacks
|
||
|
|
|
||
|
|
When you register a skill, you're teaching the shepherd about a new capability. When you contribute domain knowledge, you're sharing wisdom so future workflows avoid mistakes.
|
||
|
|
|
||
|
|
### What is a Registerable Skill?
|
||
|
|
|
||
|
|
A **skill** is a composable unit of work that:
|
||
|
|
|
||
|
|
1. **Solves a specific problem** (code analysis, deployment, notification)
|
||
|
|
2. **Has well-defined inputs/outputs** (Git repo URL → vulnerability report)
|
||
|
|
3. **Encodes domain knowledge** (retry policies, timeouts, dependencies)
|
||
|
|
4. **Scales independently** (can be retried, cached, load-balanced)
|
||
|
|
5. **Is discovered by the router** (LLM knows when/how to use it)
|
||
|
|
|
||
|
|
### Types of Registerable Knowledge
|
||
|
|
|
||
|
|
#### 1. Activity Knowledge (What can be done?)
|
||
|
|
- **Definition:** Capabilities, inputs, outputs, timeout/retry policies
|
||
|
|
- **Stored in:** `activity_knowledge_base.json`
|
||
|
|
- **Used by:** LLM router to decide *which* activities to use
|
||
|
|
|
||
|
|
```json
|
||
|
|
{
|
||
|
|
"name": "SecurityScan",
|
||
|
|
"timeout": "60s",
|
||
|
|
"maxRetries": 2,
|
||
|
|
"inputs": ["repo_url", "branch"],
|
||
|
|
"outputs": ["report", "severity_count"]
|
||
|
|
}
|
||
|
|
```
|
||
|
|
|
||
|
|
#### 2. Domain Knowledge (How should it be done?)
|
||
|
|
- **Definition:** Best practices, error patterns, optimization strategies
|
||
|
|
- **Stored in:** Semantic memory service (`poimen-memory`)
|
||
|
|
- **Used by:** LLM router to augment prompts before spec generation
|
||
|
|
|
||
|
|
```go
|
||
|
|
type MemoryContext struct {
|
||
|
|
Skills []string // Relevant skills
|
||
|
|
Lessons []string // Best practices
|
||
|
|
ErrorPatterns []string // Common failures & fixes
|
||
|
|
}
|
||
|
|
```
|
||
|
|
|
||
|
|
#### 3. Workflow Patterns (What sequences work?)
|
||
|
|
- **Definition:** Proven workflows for common scenarios
|
||
|
|
- **Stored in:** LLM prompt templates (`agent-prompts/router/AGENTS.md`)
|
||
|
|
- **Used by:** SpecBuilder to generate efficient workflows
|
||
|
|
|
||
|
|
```yaml
|
||
|
|
Patterns:
|
||
|
|
- name: "Security Review"
|
||
|
|
activities: ["CloneRepo", "SecurityScan", "GenerateReport", "NotifyStatus"]
|
||
|
|
conditions: {"branch": "main"}
|
||
|
|
```
|
||
|
|
|
||
|
|
#### 4. Skill Dependencies (What needs what?)
|
||
|
|
- **Definition:** Activity prerequisites and data flow
|
||
|
|
- **Stored in:** Activity specs + JSONPath parameter bindings
|
||
|
|
- **Used by:** Workflow validator to check graph validity
|
||
|
|
|
||
|
|
```json
|
||
|
|
{
|
||
|
|
"activity": "AnalyzeCode",
|
||
|
|
"parameters": {
|
||
|
|
"source_path": "${CloneRepo.output.repo_path}"
|
||
|
|
},
|
||
|
|
"depends_on": ["CloneRepo"]
|
||
|
|
}
|
||
|
|
```
|
||
|
|
|
||
|
|
#### 5. Execution Policies (When should it fail/retry?)
|
||
|
|
- **Definition:** Retry strategies, error handlers, fallbacks
|
||
|
|
- **Stored in:** Activity specs + catch blocks in workflows
|
||
|
|
- **Used by:** StateTransitioner to decide recovery actions
|
||
|
|
|
||
|
|
```json
|
||
|
|
{
|
||
|
|
"activity": "NotifyStatus",
|
||
|
|
"retry_policy": {
|
||
|
|
"max_attempts": 3,
|
||
|
|
"initial_interval": "1s",
|
||
|
|
"backoff_coefficient": 2.0
|
||
|
|
},
|
||
|
|
"catch_blocks": [
|
||
|
|
{
|
||
|
|
"error_type": "AuthenticationError",
|
||
|
|
"action": "log_and_continue"
|
||
|
|
}
|
||
|
|
]
|
||
|
|
}
|
||
|
|
```
|
||
|
|
|
||
|
|
### How LLM Router Uses Registered Skills
|
||
|
|
|
||
|
|
#### Flow: User Request → Knowledge-Augmented Routing
|
||
|
|
|
||
|
|
```
|
||
|
|
User: "Scan repository for security issues and deploy if safe"
|
||
|
|
↓
|
||
|
|
[1. Retrieve Memory]
|
||
|
|
Query for "security scan + deployment" patterns
|
||
|
|
Returns: Best practices, error handlers, proven sequences
|
||
|
|
↓
|
||
|
|
[2. Augment LLM Prompt]
|
||
|
|
System message includes:
|
||
|
|
- Available activities (from knowledge_base.json)
|
||
|
|
- Relevant patterns (from memory)
|
||
|
|
- Lessons learned (from memory)
|
||
|
|
↓
|
||
|
|
[3. Generate WorkflowSpec]
|
||
|
|
LLM reasons:
|
||
|
|
- "SecurityScan is first (needs fresh clone)"
|
||
|
|
- "If vulnerabilities found, stop; don't deploy"
|
||
|
|
- "Cache results per commit for 5 minutes"
|
||
|
|
↓
|
||
|
|
[4. Validate Spec]
|
||
|
|
- All activities registered? ✓
|
||
|
|
- All parameters have sources? ✓
|
||
|
|
- Timeouts reasonable? ✓
|
||
|
|
↓
|
||
|
|
[5. Execute]
|
||
|
|
RoutingWorkflow with automatic retries + error recovery
|
||
|
|
```
|
||
|
|
|
||
|
|
### Registering a New Skill (Step-by-Step)
|
||
|
|
|
||
|
|
#### Step 1: Implement the Activity
|
||
|
|
```go
|
||
|
|
// action/my_new_skill.go
|
||
|
|
package action
|
||
|
|
|
||
|
|
type MyNewSkillInput struct {
|
||
|
|
Input1 string `json:"input1"`
|
||
|
|
Input2 int `json:"input2"`
|
||
|
|
}
|
||
|
|
|
||
|
|
type MyNewSkillOutput struct {
|
||
|
|
Success bool `json:"success"`
|
||
|
|
Message string `json:"message"`
|
||
|
|
Data string `json:"data"`
|
||
|
|
}
|
||
|
|
|
||
|
|
func MyNewSkill(ctx context.Context, input *MyNewSkillInput) (*MyNewSkillOutput, error) {
|
||
|
|
// Implement logic
|
||
|
|
return &MyNewSkillOutput{
|
||
|
|
Success: true,
|
||
|
|
Message: "Executed successfully",
|
||
|
|
Data: "result...",
|
||
|
|
}, nil
|
||
|
|
}
|
||
|
|
```
|
||
|
|
|
||
|
|
#### Step 2: Add to Knowledge Base
|
||
|
|
Edit `internal/routing/activity_knowledge_base.json`:
|
||
|
|
|
||
|
|
```json
|
||
|
|
{
|
||
|
|
"name": "MyNewSkill",
|
||
|
|
"description": "Does something important",
|
||
|
|
"inputs": [
|
||
|
|
{
|
||
|
|
"name": "input1",
|
||
|
|
"type": "string",
|
||
|
|
"description": "First input parameter",
|
||
|
|
"required": true,
|
||
|
|
"examples": ["example1", "example2"]
|
||
|
|
}
|
||
|
|
],
|
||
|
|
"outputs": [
|
||
|
|
{
|
||
|
|
"name": "success",
|
||
|
|
"type": "boolean",
|
||
|
|
"description": "Whether executed successfully"
|
||
|
|
}
|
||
|
|
],
|
||
|
|
"timeout": "30s",
|
||
|
|
"maxRetries": 2,
|
||
|
|
"retryPolicy": "exponential",
|
||
|
|
"dependencies": [],
|
||
|
|
"tags": ["domain", "category"],
|
||
|
|
"knowledge": {
|
||
|
|
"bestPractices": [
|
||
|
|
"Always validate input format",
|
||
|
|
"Log all errors"
|
||
|
|
],
|
||
|
|
"errorHandling": "On ValidationError, return empty data",
|
||
|
|
"performance": "Caches results for 1 minute",
|
||
|
|
"constraints": "Requires internet connection"
|
||
|
|
}
|
||
|
|
}
|
||
|
|
```
|
||
|
|
|
||
|
|
#### Step 3: Register with Temporal Worker
|
||
|
|
```go
|
||
|
|
// cmd/worker/main.go
|
||
|
|
worker.RegisterActivity(action.MyNewSkill)
|
||
|
|
```
|
||
|
|
|
||
|
|
#### Step 4: Test Routing
|
||
|
|
```bash
|
||
|
|
go build -o worker ./cmd/worker
|
||
|
|
./worker &
|
||
|
|
|
||
|
|
./starter --route "Use MyNewSkill to process my data"
|
||
|
|
```
|
||
|
|
|
||
|
|
### Skill Registration Checklist ✅
|
||
|
|
|
||
|
|
- [ ] **Well-defined interface** — Inputs/outputs clearly typed
|
||
|
|
- [ ] **Timeout is reasonable** — Not too short (flaky), not too long (blocking)
|
||
|
|
- [ ] **Retry policy is sensible** — Network failures get retries; auth failures don't
|
||
|
|
- [ ] **Error handling is graceful** — Returns partial results, doesn't crash
|
||
|
|
- [ ] **Dependencies are clear** — Knows what other skills it needs
|
||
|
|
- [ ] **Knowledge is explicit** — Best practices documented in spec
|
||
|
|
- [ ] **Idempotent or tagged** — Safe to retry without side effects
|
||
|
|
- [ ] **Observable** — Logs important milestones
|
||
|
|
- [ ] **Tested** — Unit tests + integration tests
|
||
|
|
- [ ] **Documented** — Clear description for LLM prompts
|
||
|
|
|
||
|
|
### The Virtuous Cycle
|
||
|
|
|
||
|
|
```
|
||
|
|
1. Register Skill A
|
||
|
|
↓
|
||
|
|
2. Router uses Skill A in workflows
|
||
|
|
↓
|
||
|
|
3. Execution logs captured → lessons learned
|
||
|
|
↓
|
||
|
|
4. Register Domain Knowledge from Skill A's lessons
|
||
|
|
↓
|
||
|
|
5. Future workflows route smarter (memory-augmented)
|
||
|
|
↓
|
||
|
|
6. Register Skill B (improved by A's knowledge)
|
||
|
|
↓
|
||
|
|
7. Repeat → System gets smarter
|
||
|
|
```
|
||
|
|
|
||
|
|
**The system is self-improving.** Each addition makes it better for everyone.
|
||
|
|
|
||
|
|
---
|
||
|
|
|
||
|
|
## Project Structure
|
||
|
|
|
||
|
|
```
|
||
|
|
workflows/
|
||
|
|
├── cmd/
|
||
|
|
│ ├── worker/ # Temporal worker (registers activities & workflows)
|
||
|
|
│ └── starter/ # CLI for triggering workflows
|
||
|
|
├── action/ # Activity implementations
|
||
|
|
│ ├── analysis.go # Code analysis activity
|
||
|
|
│ ├── memory.go # Memory retrieval activity
|
||
|
|
│ ├── notification.go # Status notifications
|
||
|
|
│ ├── logger.go # Structured logging
|
||
|
|
│ └── router.go # Router activity orchestration
|
||
|
|
├── internal/routing/
|
||
|
|
│ ├── types.go # WorkflowSpec, ActivitySpec types
|
||
|
|
│ ├── knowledge_base.go # Activity knowledge base loader
|
||
|
|
│ ├── llm_client.go # LLM provider interface
|
||
|
|
│ ├── llm_router.go # Natural language → WorkflowSpec
|
||
|
|
│ ├── provider.go # LLMProvider implementations (caching, retry)
|
||
|
|
│ ├── spec_builder.go # SpecBuilder interface & implementations
|
||
|
|
│ ├── executor.go # Validators, executors, state transitioners
|
||
|
|
│ ├── activity_knowledge_base.json # Activity registry
|
||
|
|
│ └── *_test.go # Unit tests
|
||
|
|
├── statemachine/
|
||
|
|
│ └── routing_workflow.go # RoutingWorkflow state machine executor
|
||
|
|
├── agent-prompts/
|
||
|
|
│ └── router/AGENTS.md # LLM prompts for routing decisions
|
||
|
|
├── k8s/
|
||
|
|
│ ├── worker-deployment.yaml
|
||
|
|
│ └── configmap.yaml
|
||
|
|
├── examples/
|
||
|
|
│ ├── service_integration.go
|
||
|
|
│ └── wait_for_task_test.go
|
||
|
|
├── tests/
|
||
|
|
│ └── *_test.go # Integration & E2E tests
|
||
|
|
├── docs/
|
||
|
|
│ └── *.md # Architecture & design docs
|
||
|
|
└── README.md # This file
|
||
|
|
```
|
||
|
|
|
||
|
|
---
|
||
|
|
|
||
|
|
## Configuration
|
||
|
|
|
||
|
|
### Environment Variables
|
||
|
|
```bash
|
||
|
|
# LLM Configuration
|
||
|
|
LLM_ENDPOINT="https://api.riotpiao.com/v1/chat/completions"
|
||
|
|
LLM_MODEL="reasoning"
|
||
|
|
LLM_API_KEY="your-api-key"
|
||
|
|
|
||
|
|
# Temporal Configuration
|
||
|
|
TEMPORAL_HOST_URL="localhost:7233"
|
||
|
|
TEMPORAL_NAMESPACE="default"
|
||
|
|
|
||
|
|
# Memory Service
|
||
|
|
MEMORY_SERVICE_URL="http://poimen-memory.poimen.svc.cluster.local:8080"
|
||
|
|
|
||
|
|
# Logging
|
||
|
|
LOG_LEVEL="info"
|
||
|
|
```
|
||
|
|
|
||
|
|
### LLMRouter Configuration (Programmatic)
|
||
|
|
```go
|
||
|
|
config := &routing.LLMRouterConfig{
|
||
|
|
KnowledgeBase: kb,
|
||
|
|
LLMProviders: []routing.LLMProvider{
|
||
|
|
openai.NewOpenAIProvider("gpt-4"),
|
||
|
|
claude.NewClaudeProvider("claude-3"),
|
||
|
|
local.NewLocalProvider(":8000"),
|
||
|
|
},
|
||
|
|
SpecBuilders: []routing.SpecBuilder{
|
||
|
|
routing.NewDefaultSpecBuilder(),
|
||
|
|
routing.NewCronSpecBuilder(),
|
||
|
|
},
|
||
|
|
Validators: []routing.WorkflowValidator{
|
||
|
|
routing.NewStateGraphValidator(),
|
||
|
|
routing.NewActivityAvailabilityValidator(),
|
||
|
|
},
|
||
|
|
CacheProvider: routing.NewCachingLLMProvider(cache),
|
||
|
|
}
|
||
|
|
router := routing.NewLLMRouter(config)
|
||
|
|
```
|
||
|
|
|
||
|
|
---
|
||
|
|
|
||
|
|
## Testing
|
||
|
|
|
||
|
|
### Unit Tests
|
||
|
|
```bash
|
||
|
|
go test ./... -v
|
||
|
|
```
|
||
|
|
|
||
|
|
### Coverage Report
|
||
|
|
```bash
|
||
|
|
go test ./... -cover -coverprofile=coverage.out
|
||
|
|
go tool cover -html=coverage.out
|
||
|
|
```
|
||
|
|
|
||
|
|
### Integration Tests
|
||
|
|
```bash
|
||
|
|
# Requires Temporal running
|
||
|
|
go test -tags=integration ./tests/...
|
||
|
|
```
|
||
|
|
|
||
|
|
### Current Coverage
|
||
|
|
- `internal/routing`: 45.9%
|
||
|
|
- `action`: ~60%
|
||
|
|
- `statemachine`: ~50%
|
||
|
|
|
||
|
|
See [`poimen-docs/CRAP_ANALYSIS.md`](../poimen-docs/CRAP_ANALYSIS.md) for complexity metrics.
|
||
|
|
|
||
|
|
---
|
||
|
|
|
||
|
|
## CRAP Score Improvements
|
||
|
|
|
||
|
|
Refactoring focused on reducing cyclomatic complexity via helper extraction and strategy patterns:
|
||
|
|
|
||
|
|
| Function | Before | After | Improvement |
|
||
|
|
|----------|--------|-------|-------------|
|
||
|
|
| `buildCronSpec` | 156 | 5 | 97% ↓ |
|
||
|
|
| `buildParameters` | 90 | 6 | 93% ↓ |
|
||
|
|
| `RoutingWorkflow` | 132 | 7 | 95% ↓ |
|
||
|
|
| `DeploymentPreCheckActivity` | 110 | 6 | 95% ↓ |
|
||
|
|
|
||
|
|
**Target:** CRAP < 10 for all production functions ✅
|
||
|
|
|
||
|
|
---
|
||
|
|
|
||
|
|
## CI/CD Pipeline
|
||
|
|
|
||
|
|
### GitHub Actions
|
||
|
|
- **On Push:** Run tests, lint, coverage checks
|
||
|
|
- **On PR:** Automated tests + code review gates
|
||
|
|
- **On Release:** Build binaries for macOS/Linux/Windows, publish to registry
|
||
|
|
|
||
|
|
```yaml
|
||
|
|
# Typical workflow
|
||
|
|
1. Unit tests (all packages)
|
||
|
|
2. Go vet & lint checks
|
||
|
|
3. Integration tests (Temporal required)
|
||
|
|
4. Coverage validation (>40% threshold)
|
||
|
|
5. Deploy staging worker (if main branch)
|
||
|
|
```
|
||
|
|
|
||
|
|
### Build & Deploy
|
||
|
|
```bash
|
||
|
|
# Local build
|
||
|
|
make build # Builds cmd/worker and cmd/starter
|
||
|
|
|
||
|
|
# Docker build (K8s deployment)
|
||
|
|
docker build -t your-registry/poimen-worker:latest .
|
||
|
|
|
||
|
|
# Deploy to K8s
|
||
|
|
kubectl apply -f k8s/worker-deployment.yaml
|
||
|
|
```
|
||
|
|
|
||
|
|
---
|
||
|
|
|
||
|
|
## Architecture Decisions
|
||
|
|
|
||
|
|
### Why Temporal?
|
||
|
|
- **Durable Execution:** Survives crashes, network splits, process restarts
|
||
|
|
- **Visibility:** Full workflow history in UI + API
|
||
|
|
- **Failure Handling:** Built-in retry, timeout, compensation logic
|
||
|
|
- **Scalability:** Millions of concurrent workflows
|
||
|
|
|
||
|
|
### Why LLM Router?
|
||
|
|
- **Intent Understanding:** Captures user intent without rigid form parsing
|
||
|
|
- **Adaptive Workflows:** LLM optimizes step ordering based on knowledge base
|
||
|
|
- **Extensibility:** New activities auto-discovered; no hardcoding
|
||
|
|
- **Memory Augmentation:** Context-aware decisions via semantic search
|
||
|
|
|
||
|
|
### Why Generic State Machine?
|
||
|
|
- **Composability:** Any activity becomes a workflow step
|
||
|
|
- **Reusability:** One state machine handles all workflow patterns
|
||
|
|
- **Testability:** Validates JSON specs independently of execution
|
||
|
|
- **Debuggability:** JSONPath-based tracing for data flow
|
||
|
|
|
||
|
|
### Plugin Architecture Benefits
|
||
|
|
- **Provider Swapping:** Switch LLM backends without code changes
|
||
|
|
- **Strategy Injection:** Multiple spec builders, validators in pipeline
|
||
|
|
- **Decorator Pattern:** Caching, retries, logging as independent layers
|
||
|
|
- **Backward Compatibility:** `NewLLMRouterDefault()` maintains old API
|
||
|
|
|
||
|
|
---
|
||
|
|
|
||
|
|
## Documentation
|
||
|
|
|
||
|
|
- **[Routing Workflow Spec](./docs/ROUTING_WORKFLOW_SPEC.md)** — Complete spec format reference
|
||
|
|
- **[Activity Knowledge Base](./docs/ACTIVITY_KNOWLEDGE_BASE.md)** — Activity definitions & capabilities
|
||
|
|
- **[LLM Router Prompts](./agent-prompts/router/AGENTS.md)** — System prompts for routing decisions
|
||
|
|
- **[CRAP Analysis](../poimen-docs/CRAP_ANALYSIS.md)** — Code complexity refactoring details
|
||
|
|
- **[Implementation Notes](./README_IMPLEMENTATION.md)** — Design decisions & trade-offs
|
||
|
|
|
||
|
|
---
|
||
|
|
|
||
|
|
## Contributing
|
||
|
|
|
||
|
|
### Adding a New Activity
|
||
|
|
1. Implement `Activity` interface in `action/`
|
||
|
|
2. Add spec to `internal/routing/activity_knowledge_base.json`
|
||
|
|
3. Register in `cmd/worker/main.go`
|
||
|
|
4. Test with CLI: `./starter --route "..."`
|
||
|
|
|
||
|
|
### Extending the Router
|
||
|
|
1. Implement `LLMProvider`, `SpecBuilder`, or `WorkflowValidator` interface
|
||
|
|
2. Register in `LLMRouterConfig`
|
||
|
|
3. Add unit tests in `internal/routing/*_test.go`
|
||
|
|
4. Update docs
|
||
|
|
|
||
|
|
### Code Style
|
||
|
|
- Run `gofmt` and `go vet` before commit
|
||
|
|
- Write tests for all exported functions
|
||
|
|
- Keep CRAP score < 10
|
||
|
|
- Document non-obvious logic
|
||
|
|
|
||
|
|
---
|
||
|
|
|
||
|
|
## Troubleshooting
|
||
|
|
|
||
|
|
### "Activity not registered"
|
||
|
|
- Ensure `cmd/worker` is running
|
||
|
|
- Check `activity_knowledge_base.json` includes activity name
|
||
|
|
- Verify activity name in workflow spec matches registration
|
||
|
|
|
||
|
|
### "Workflow failed to execute"
|
||
|
|
- Check Temporal logs: `temporal workflow show -w <workflow-id>`
|
||
|
|
- Review activity output in Temporal UI (`:8081`)
|
||
|
|
- Enable debug logging: `LOG_LEVEL=debug`
|
||
|
|
|
||
|
|
### "LLM router timeout"
|
||
|
|
- Increase `LLMRequestTimeout` in config
|
||
|
|
- Check API endpoint reachability: `curl https://api.riotpiao.com/v1/chat/completions`
|
||
|
|
- Review LLM prompt length (may be too complex)
|
||
|
|
|
||
|
|
### "Memory service not responding"
|
||
|
|
- Verify `MEMORY_SERVICE_URL` endpoint
|
||
|
|
- Check K8s pod: `kubectl logs -l app=poimen-memory`
|
||
|
|
- Graceful fallback: RetrieveMemoryActivity is optional
|
||
|
|
|
||
|
|
---
|
||
|
|
|
||
|
|
## Roadmap
|
||
|
|
|
||
|
|
- [ ] **Real Cron Scheduling** — Implement Temporal ScheduleClient for recurring workflows
|
||
|
|
- [ ] **Additional LLM Providers** — OpenAI, Claude, local model integrations
|
||
|
|
- [ ] **Provider Auto-Discovery** — Config-driven provider registration
|
||
|
|
- [ ] **Improved Coverage** — Target 60%+ in all packages
|
||
|
|
- [ ] **Web Dashboard** — Real-time workflow execution dashboard
|
||
|
|
- [ ] **Approval Workflows** — Multi-stage human approval gates
|
||
|
|
|
||
|
|
---
|
||
|
|
|
||
|
|
## License
|
||
|
|
|
||
|
|
[Your License Here]
|
||
|
|
|
||
|
|
## Support
|
||
|
|
|
||
|
|
- **Issues:** GitHub Issues
|
||
|
|
- **Discussions:** GitHub Discussions
|
||
|
|
- **Security:** See SECURITY.md
|
||
|
|
|
||
|
|
---
|
||
|
|
|
||
|
|
**Built with ❤️ using Temporal, Go, and LLMs.**
|