docs: temporal + graph RAG integration with unified query

This commit is contained in:
Test
2026-09-05 05:45:21 -07:00
parent 9624f0e18d
commit a461e9799a
4 changed files with 1727 additions and 2 deletions
File diff suppressed because it is too large Load Diff
+712
View File
@@ -0,0 +1,712 @@
# Workflows + Graph RAG Integration
Align Poimen Workflows with Poimen Memory System API for versioned relations, semantic queries, and intelligent canvas reasoning.
**Key Features:**
- Versioned workflow relations following `/memory/entities/{id}/versions` pattern
- Semantic edge queries via `/workflows/{id}/query/semantic/edges`
- Relation wording as Facts (matching Memory System edges schema)
- Point-in-time canvas reconstruction via `?as_of=timestamp`
- Ranking profiles for relation importance
- Automatic Graph RAG indexing
---
## Architecture
```
┌─────────────────────────────────┐
│ Workflow Canvas (React Flow) │
│ - Nodes (activities) │
│ - Edges (connections) │
└──────────────┬──────────────────┘
│ PUT /workflows/{id}
┌─────────────────────────────────┐
│ CanvasReasonerActivity │
│ - Suggest edges │
│ - Validate compatibility │
│ - Generate change reasoning │
└──────────────┬──────────────────┘
│ suggested_edges + reasoning
┌─────────────────────────────────┐
│ Workflows API Server │
│ - Update canvas in DB │
│ - Create version entry │
│ - Store change metadata │
└──────────────┬──────────────────┘
│ canvas_version, relations
┌─────────────────────────────────┐
│ Graph RAG Backend │
│ - Store versioned relations │
│ - Index relation wording │
│ - Enable semantic queries │
└─────────────────────────────────┘
```
---
## Data Model
### Workflow Canvas Version
```sql
-- In memory.workflows_versions (new table)
CREATE TABLE workflow_versions (
id UUID PRIMARY KEY,
workflow_id UUID NOT NULL REFERENCES workflows(id),
customer_id UUID NOT NULL,
version INT NOT NULL,
canvas JSONB NOT NULL, -- {nodes, edges}
changed_by UUID,
change_reason TEXT,
change_type VARCHAR(50), -- 'manual', 'auto_reasoned', 'import'
reasoner_confidence FLOAT,
reasoning_metadata JSONB, -- LLM reasoning output
created_at TIMESTAMP DEFAULT NOW(),
UNIQUE(workflow_id, version),
FOREIGN KEY(workflow_id, customer_id)
REFERENCES workflows(id, customer_id)
);
-- In memory.workflow_relations (replaces simple edges)
CREATE TABLE workflow_relations (
id UUID PRIMARY KEY,
workflow_id UUID NOT NULL,
version INT NOT NULL,
source_node_id VARCHAR(255),
target_node_id VARCHAR(255),
relation_type VARCHAR(100), -- 'data-flow', 'dependency', 'conditional', 'parallel'
relation_label TEXT, -- Human-readable: "SecurityScan outputs issues → Report inputs requirements"
relation_wording JSONB, -- {verb, object, context}
metadata JSONB, -- {source_output_type, target_input_type, compatibility_score}
created_at TIMESTAMP,
FOREIGN KEY(workflow_id, version)
REFERENCES workflow_versions(id, version),
INDEX (workflow_id, version)
);
-- In memory.relation_changes (for Graph RAG indexing)
CREATE TABLE relation_changes (
id UUID PRIMARY KEY,
workflow_id UUID,
version_from INT,
version_to INT,
change_type VARCHAR(50), -- 'added', 'removed', 'modified'
relation_id UUID REFERENCES workflow_relations(id),
source_node_id VARCHAR(255),
target_node_id VARCHAR(255),
old_wording JSONB,
new_wording JSONB,
change_reason TEXT,
change_timestamp TIMESTAMP,
reasoner_confidence FLOAT,
INDEX (workflow_id, version_to)
);
```
### Relation Wording Schema
```json
{
"id": "edge-1",
"source": "clone-repo-1",
"target": "analyze-code-1",
"relation_type": "data-flow",
"relation_label": "CloneRepo outputs path → AnalyzeCode requires path",
"relation_wording": {
"verb": "outputs",
"source_output": "path (string): Local filesystem path where repo was cloned",
"target_input": "path (string, required): Local filesystem path to analyze",
"connection_type": "direct-map",
"confidence": 0.98,
"notes": "Perfect type match between CloneRepo.path and AnalyzeCode.path"
},
"metadata": {
"source_activity": "CloneRepoActivity",
"target_activity": "AnalyzeCodeActivity",
"output_type": "string",
"input_type": "string",
"compatibility_score": 0.98,
"requires_transformation": false,
"semantic_match": "File path passes directly"
},
"change_history": [
{
"version": 2,
"action": "added",
"reason": "LLM reasoner suggested data-flow connection",
"confidence": 0.98,
"timestamp": "2025-09-05T10:00:00Z"
}
]
}
```
---
## Unified API Endpoints
All workflow queries follow the unified endpoint pattern from Memory System.
### 1. Unified Workflow Query
**Endpoint:** `POST /workflows/{id}/query`
This is the primary endpoint for all workflow canvas queries (replaces separate search endpoints).
**Request:**
```json
{
"query": "how does code analysis flow into security scanning",
"search_type": "edges|entities|all",
"version": 3,
"relation_type": "data-flow",
"confidence_floor": 0.7,
"top_k": 10,
"find_paths": true,
"target_node_id": "security-scan-1",
"max_path_depth": 3,
"ranking_profile": "default",
"include_reasoning": true
}
```
**Response (200 OK):**
```json
{
"workflow_id": "workflow-1",
"query": "how does code analysis flow into security scanning",
"search_type": "edges",
"version": 3,
"execution_time_ms": 145,
"results": [
{
"id": "edge_analyze_scan",
"source_node_id": "analyze-code-1",
"source_name": "AnalyzeCodeActivity",
"target_node_id": "security-scan-1",
"target_name": "SecurityScanActivity",
"relation_type": "data-flow",
"relation_label": "AnalyzeCode outputs metrics → SecurityScan requires code structure",
"relation_wording": {
"verb": "provides-input-for",
"source_output": "metrics (object): Code quality and structural metrics",
"target_input": "path (string): Directory to scan",
"connection_type": "requires-transformer",
"confidence": 0.85,
"semantic_match": "Analysis metrics can guide security scan prioritization"
},
"similarity_score": 0.92,
"confidence": 0.85,
"created_at": "2025-09-05T10:00:00Z",
"metadata": {
"source": "canvas://workflow-1:v3",
"tags": ["code-review", "security"]
}
}
],
"paths": [
{
"source_id": "analyze-code-1",
"target_id": "security-scan-1",
"path_count": 1,
"shortest_distance": 1,
"paths_found": [
{
"node_ids": ["analyze-code-1", "security-scan-1"],
"relation_types": ["data-flow"],
"distance": 1,
"total_confidence": 0.85
}
]
}
],
"total_count": 1,
"has_more": false,
"ranking_profile": "default"
}
```
---
### 2. Update Workflow Canvas (with Versioning)
**Endpoint:** `PUT /workflows/{id}`
**Request:**
```json
{
"nodes": [...],
"edges": [...],
"auto_reason": true,
"change_reason": "User connected CloneRepo to AnalyzeCode",
"user_id": "uuid"
}
```
**Response:**
```json
{
"id": "workflow-1",
"version": 3,
"canvas": {
"nodes": [...],
"edges": [...]
},
"version_info": {
"version_number": 3,
"created_at": "2025-09-05T10:05:00Z",
"created_by": "user-uuid",
"change_reason": "User connected CloneRepo to AnalyzeCode",
"change_type": "manual"
},
"relation_updates": {
"added": [
{
"source": "clone-repo-1",
"target": "analyze-code-1",
"relation_wording": {
"verb": "outputs",
"source_output": "path: Local filesystem path where repo was cloned",
"target_input": "path (required): Local filesystem path to analyze",
"confidence": 0.98
}
}
],
"removed": [],
"modified": []
}
}
```
---
### 2. Get Relation Version History (Memory System Pattern)
**Endpoint:** `GET /workflows/{id}/relations/{edge_id}/versions`
Follows `/memory/entities/{id}/versions` pattern from Memory System.
**Response:**
```json
{
"edge_id": "edge_1",
"workflow_id": "workflow-1",
"source": "clone-repo-1",
"target": "analyze-code-1",
"versions": [
{
"version_num": 1,
"operation": "CREATE",
"snapshot": {
"relation_type": "data-flow",
"relation_label": "CloneRepo → AnalyzeCode",
"relation_wording": {
"verb": "connects-to",
"confidence": 0.75
}
},
"changed_at": "2025-09-04T12:00:00Z",
"changed_by": "system",
"fields_changed": ["relation_type", "relation_wording"]
},
{
"version_num": 2,
"operation": "UPDATE",
"snapshot": {
"relation_type": "data-flow",
"relation_label": "CloneRepo outputs path → AnalyzeCode requires path",
"relation_wording": {
"verb": "outputs",
"source_output": "path (string)",
"target_input": "path (string, required)",
"confidence": 0.98
}
},
"changed_at": "2025-09-05T10:00:00Z",
"changed_by": "reasoner-activity",
"fields_changed": ["relation_wording", "relation_label"]
}
],
"total_versions": 2,
"current_version": 2
}
```
---
### 3. Edge Diff (Versioning Pattern)
**Endpoint:** `POST /workflows/{id}/relations/diff`
Follows `/memory/entities/diff` pattern.
**Request:**
```json
{
"edge_id": "edge_1",
"from_version": 1,
"to_version": 2
}
```
**Response:**
```json
{
"edge_id": "edge_1",
"from_version": 1,
"to_version": 2,
"source": "clone-repo-1",
"target": "analyze-code-1",
"diff": {
"added_fields": {},
"removed_fields": {},
"modified_fields": {
"relation_wording": {
"old": {
"verb": "connects-to",
"confidence": 0.75
},
"new": {
"verb": "outputs",
"source_output": "path (string)",
"target_input": "path (string, required)",
"confidence": 0.98
}
}
}
},
"change_timeline": [
{
"version": 1,
"confidence": 0.75,
"changed_at": "2025-09-04T12:00:00Z"
},
{
"version": 2,
"confidence": 0.98,
"changed_at": "2025-09-05T10:00:00Z"
}
],
"editors_involved": ["system", "reasoner-activity"]
}
```
---
### 4. Point-in-Time Canvas (Versioning Pattern)
**Endpoint:** `GET /workflows/{id}?at_version={v}` or `?as_of=2025-09-05T10:00:00Z`
Follows `/memory/entities/at` pattern.
**Response:**
```json
{
"workflow_id": "workflow-1",
"version": 2,
"as_of_timestamp": "2025-09-05T10:00:00Z",
"canvas": {
"nodes": [...],
"edges": [...]
},
"relations": [
{
"id": "edge_1",
"source": "clone-repo-1",
"target": "analyze-code-1",
"relation_type": "data-flow",
"relation_label": "CloneRepo outputs path → AnalyzeCode requires path",
"relation_wording": {
"verb": "outputs",
"source_output": "path (string)",
"target_input": "path (string, required)",
"confidence": 0.98
},
"version": 2
}
],
"metadata": {
"version_number": 2,
"created_at": "2025-09-05T10:00:00Z",
"changed_by": "reasoner-activity",
"change_reason": "LLM reasoner refined relation wording"
}
}
```
---
### 5. Bulk Import Canvas (with Relations)
**Endpoint:** `POST /workflows/{id}/import`
**Request:**
```json
{
"canvas": {
"nodes": [...],
"edges": [...]
},
"relations": [
{
"source": "n1",
"target": "n2",
"relation_type": "data-flow",
"relation_wording": {
"verb": "outputs",
"source_output": "result (string)",
"target_input": "input (string, required)"
}
}
],
"change_reason": "Imported from external workflow system"
}
```
**Response:**
```json
{
"workflow_id": "workflow-1",
"version": 4,
"canvas": {...},
"relations": {...},
"import_metadata": {
"imported_nodes": 5,
"imported_edges": 4,
"validation_status": "success",
"indexing_status": "queued_for_graph_rag"
}
}
```
---
## Integration with CanvasReasonerActivity
### Flow
```
User edits canvas
PUT /workflows/{id} with auto_reason=true
Backend calls CanvasReasonerActivity
LLM suggests edges + reasoning
Generate relation_wording from suggestions
Create workflow_version entry
Create workflow_relations entries
Index relations in Graph RAG
Response includes suggested_edges + relation_wording
```
### Response Structure
```json
{
"version": 3,
"suggested_edges": [
{
"source": "clone-1",
"target": "analyze-1"
}
],
"reasoning": "Standard code review workflow",
"confidence": 0.92,
"relation_wordings": [
{
"source": "clone-1",
"target": "analyze-1",
"relation_wording": {
"verb": "outputs",
"source_output": "path (string): Cloned repository path",
"target_input": "path (string, required): Directory to analyze",
"connection_type": "direct-map",
"confidence": 0.98,
"semantic_description": "Repository path flows from clone operation to analysis"
}
}
]
}
```
---
## Graph RAG Indexing
### Relations Indexed
Each workflow relation creates Graph RAG entities:
```
Node: {
id: "clone-repo-1",
type: "activity",
name: "CloneRepoActivity",
workflow_id: "workflow-1",
version: 3
}
Node: {
id: "analyze-code-1",
type: "activity",
...
}
Edge: {
id: "rel-1",
source: "clone-repo-1",
target: "analyze-code-1",
type: "data-flow",
label: "outputs path",
wording: {...},
version: 3,
created_at: "2025-09-05T10:00:00Z"
}
```
### Semantic Queries
Users can query like:
- *"Which activities receive data from CloneRepo?"*
- *"What's the data flow from Analyze to Report?"*
- *"Which relations were added in version 3?"*
- *"Show me all type-mismatched connections"*
- *"Find workflows with Security Scan that require approval"*
---
## Versioning Strategy
### Version Numbering
- Increment on every canvas change
- Track change_type: `manual`, `auto_reasoned`, `import`, `rag_query_applied`
- Store reasoner_confidence for auto changes
### Relation Wording Versions
- Each relation has independent wording history
- Confidence scores tracked per version
- Sources: user input, LLM reasoner, import, RAG query
### Changelog
```json
{
"workflow_id": "workflow-1",
"total_versions": 5,
"changes": [
{
"version": 1,
"type": "created",
"timestamp": "2025-09-04T10:00:00Z",
"user": "system",
"reason": "Workflow initialized"
},
{
"version": 2,
"type": "auto_reasoned",
"timestamp": "2025-09-04T12:00:00Z",
"reasoner_confidence": 0.89,
"changes": {
"edges_added": 4,
"edges_modified": 0
}
},
{
"version": 3,
"type": "manual",
"timestamp": "2025-09-05T10:05:00Z",
"user": "user-uuid",
"reason": "Connected SecurityScan to Report manually"
}
]
}
```
---
## Data Flow Example
### Scenario: User drops SecurityScan node, system suggests connection
1. **User Action:** Drops SecurityScan node onto existing workflow
2. **Frontend Call:** `PUT /workflows/{id}` with new nodes + `auto_reason=true`
3. **Backend:**
- Saves canvas to `workflow_versions` v.3
- Calls CanvasReasonerActivity
4. **LLM Reasoning:**
- Analyzes: AnalyzeCode (outputs: quality, metrics) → SecurityScan (inputs: path, depth)
- Suggests: Add transformer node OR use metrics for decision
- Confidence: 0.85 (type mismatch, requires transformation)
5. **Relation Wording:**
```json
{
"source": "analyze-code-1",
"target": "security-scan-1",
"relation_wording": {
"verb": "provides-context-for",
"source_output": "quality (object): Code quality metrics",
"target_input": "path (string): Directory to scan",
"connection_type": "requires-transformer",
"reasoning": "Metrics inform which files to prioritize in scanning"
}
}
```
6. **Graph RAG:** Relations indexed automatically
7. **Response:** Frontend shows suggestions with natural language descriptions
---
## Implementation Checklist
- [ ] Create `workflow_versions` table in memory DB
- [ ] Create `workflow_relations` table with wording schema
- [ ] Create `relation_changes` table for tracking modifications
- [ ] Update CanvasReasonerActivity to generate relation wordings
- [ ] Implement versioned CRUD endpoints (PUT, GET, DIFF)
- [ ] Add Graph RAG indexing on relation creation
- [ ] Implement semantic query endpoint
- [ ] Add changelog view
- [ ] Test point-in-time reconstruction
- [ ] Document API in homelab-frontend/API.md
---
## Relation Wording Language
### Verbs (relation_type → verb mapping)
| Type | Verbs | Example |
|------|-------|---------|
| `data-flow` | outputs, inputs, receives, provides | "CloneRepo outputs path → AnalyzeCode inputs path" |
| `dependency` | must-complete-before, depends-on, requires | "SecurityScan depends-on AnalyzeCode completion" |
| `conditional` | triggers-if, branches-on, routes-to | "ApproveWorkflow branches-on result approval" |
| `parallel` | runs-alongside, concurrent-with, independent-of | "Notify runs-alongside Report generation" |
| `transformation` | transforms, converts, maps, adapts | "LLMTransform converts metrics → deployment plan" |
### Confidence Scoring
- **0.9-1.0:** Perfect match (type-compatible, direct data flow)
- **0.7-0.9:** Good match (semantic fit, minor transformation needed)
- **0.5-0.7:** Possible match (requires user confirmation)
- **<0.5:** Poor match (suggest removal or transformer)
+1 -1
View File
@@ -9,6 +9,6 @@ metadata:
app.kubernetes.io/name: poimen app.kubernetes.io/name: poimen
app.kubernetes.io/component: orchestrator app.kubernetes.io/component: orchestrator
data: data:
GIT_COMMIT: "8cfa23e5d" # Updated automatically by CI/CD GIT_COMMIT: "64fa03cae" # Updated automatically by CI/CD
GIT_BRANCH: "main" GIT_BRANCH: "main"
DEPLOYMENT_DATE: "2026-09-05" DEPLOYMENT_DATE: "2026-09-05"
+1 -1
View File
@@ -13,7 +13,7 @@ spec:
labels: labels:
app: poimen-worker app: poimen-worker
annotations: annotations:
git-commit: "8cfa23e5d" # ✅ Updated on each push, triggers rolling restart git-commit: "64fa03cae" # ✅ Updated on each push, triggers rolling restart
deployment-date: "2026-09-05" deployment-date: "2026-09-05"
spec: spec:
containers: containers: