Files
Test 87ceea3d30 feat(T2.4): implement fast lessons file indexing
- Add internal/indexing package for lessons index
- Implement LessonIndex with multi-field index structure
- Index by task type, activity type, failure type, and pattern
- Fast lookups: O(1) map access for all query types
- Build from JSONL file with streaming parse
- Support incremental lesson addition
- Query operations with optional AND logic
- Time range queries for temporal analysis
- Similarity search by failure message substring
- Most frequent failures ranking
- 20 indexing tests, all passing

Features:
- FindByTaskType() - query by task type
- FindByActivityType() - query by activity type
- FindByFailureType() - query by failure type
- FindByPattern() - query by pattern
- FindSimilar() - substring search in failure messages
- QueryMultiple() - AND logic for multi-field queries
- GetByTimeRange() - temporal range queries
- GetMostFrequentFailures() - ranked by frequency
- BuildFromFile() - load from JSONL
- AddLesson() - incremental updates

Performance Verified:
- Lookup < 10ms for 1000s entries ✓
- <10ms for 10,000 entries ✓
- Concurrent queries supported ✓
- O(1) average lookup complexity
- Index rebuilding efficient

Test Coverage:
- 20 indexing tests (build, query, range, stats)
- Latency verification (< 10ms)
- Concurrency testing
- Time range queries
- Multi-field queries
- Large dataset support (10k entries)

Index Structures:
- lessons: ID -> Lesson (full lookup)
- byTaskType: TaskType -> []*Lesson
- byActivityType: ActivityType -> []*Lesson
- byFailureType: FailureType -> []*Lesson
- byPattern: Pattern -> []*Lesson
- All RWMutex-protected for thread safety

Next: T2.5 (Git operation batching)
2026-08-23 17:21:23 -07:00

1.8 KiB
Raw Permalink Blame History

Task Board — Milestone T2: Scale & Performance

Submilestone: T2 (Distributed execution, caching, performance optimization)

ID Scope Status Branch Verification
T2.1 Activity result caching: deduplicate repeated LLM calls for same task state [x] task/T2.1 Implementer called 2x on same code → second call returns cached Implementer output
T2.2 Parallel task dispatch: multiple T0.x tasks execute truly concurrently (not sequential) [x] task/T2.2 9 tasks complete in ~1/9 total time (wall-clock speedup measured)
T2.3 Prompt template caching: pre-compile Go templates on worker startup [x] task/T2.3 Template render latency < 100ms (vs parse+render each time)
T2.4 Lessons file indexing: fast lookup of past failures without full file scan [x] task/T2.4 Query lessons by task type → return in < 10ms for 1000s of entries
T2.5 Git operation batching: combine multiple worktree commits into single push/merge [ ] task/T2.5 N tasks → 1 push (vs N pushes), measured via git ref-log
T2.6 LLM request batching: group similar Implementer calls into one API request [ ] task/T2.6 3 implementer tasks → 1 Anthropic API call with batch input (vs 3 separate calls)
T2.7 Workflow history pruning: trim old task unit outputs from orchestrator history [ ] task/T2.7 Continue-as-new cycle history size constant despite 1000s of task units completed
T2.8 Distributed lock optimization: replace flock with Redis/etcd for multi-pod scenarios [ ] task/T2.8 5 concurrent orchestrators on different pods share FS safely via distributed lock

Submission Criteria

All T2.1T2.8 marked [x] → submilestone complete → squash-merge task/T2.* to main.