## Fixes
- [#3] Status events now add/update a status ActivityNode in timeline (upsertWsStatusNode)
- [#4] Done callback adds done node; error callback adds error node + sets session.status="error"
- [#5] handleRegenerate fully synced with workspace: status/tool/card/done/error all handled
- [#1][#2] localStorage persistence: completed/error sessions auto-saved, lazy-loaded on demand
- [#1] handleSelectConversation preloads workspace sessions from localStorage for history messages
- Refactored completeWsSession to include done ActivityNode in timeline
- Added errorWsSession, upsertWsStatusNode, saveWsToStorage, loadWsFromStorage helpers
## Known limitation
- [#7] workspace_card merge:true not yet used by backend (all cards are append-only for now)
- History workspace recovery depends on localStorage (browser-local, not cross-device)
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
- Add _extract_output_str() to unwrap MCP ToolMessage content objects
instead of calling str() on the raw object
- Handle read_url url param when LLM passes a list instead of a string
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
- Add langchain-community dependency for GoogleSerperAPIWrapper
- Add serper_api_key to config with default key
- Create app/tools/serper.py with async serper_search tool
- Register "serper" key in tools/__init__.py (independent from "search"/Jina MCP)
- Add tool title and input/output summaries in chat.py
- Set SERPER_API_KEY on Azure App Service (Operation resource group)
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
- Add has_tool_activity flag to distinguish first agent pass (thinking)
from post-tool agent pass (generating)
- Replace on_chain_start debug logging with status SSE emission
- Filter to only graph-level agent nodes via graph:step: tag prefix
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
- tool_start/tool_end/tool_error SSE payloads now include call_id field
derived from LangGraph run_id for reliable tool event correlation
- tool_start_ts dict keyed by call_id instead of tool_name to handle
concurrent calls to the same tool
- Added on_chain_start debug logging to observe chain names and metadata
(no SSE emission yet, observation only)
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
Add CoT execution trace to the SSE stream: status events for stage
transitions, enriched tool_start with input_summary, tool_end with
output_summary and duration_ms, and tool_error for failed tool calls.
Helper functions _sse, _summarize_input, _summarize_output provide
human-readable summaries for each tool type.
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
When the frontend passes a conversation_id that does not exist in the
conversations table (e.g. new conversation before first message, or
invalid ID), the INSERT into attachments failed with
ForeignKeyViolationError.
Now upload_attachment checks session.get(Conversation, conversation_id)
before inserting. If the conversation does not exist, conversation_id is
set to None and the attachment is stored as unlinked (blob path uses
_unlinked/ prefix). This avoids the FK constraint error while keeping
the file safely uploaded.
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
When a tool call crashes (e.g. httpx timeout in kb_search), the LangGraph
checkpoint retains an AIMessage with tool_calls but no corresponding
ToolMessage. Subsequent requests to the same conversation_id fail with:
ValueError: Found AIMessages with tool_calls that do not have a
corresponding ToolMessage
Now the except block in _stream_response detects this specific ValueError
by checking for "tool_calls" and "ToolMessage" in the error string, then
calls checkpointer.adelete_thread() to purge the corrupted thread state.
The frontend receives {"type":"error","content":"对话状态异常,已自动重置..."}
followed by {"type":"done"}, so the user can simply resend their message.
API confirmed: AsyncPostgresSaver.adelete_thread(thread_id) deletes from
checkpoints, checkpoint_blobs, and checkpoint_writes tables for the thread.
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
KB Agent runs on Azure App Service which cold-starts after idle periods,
taking 20-30s to respond. The previous 15s timeout caused httpx.ReadTimeout
that crashed the ReAct agent tool node and killed the SSE stream.
Changes:
- Increase kb_agent_search_timeout_sec default from 15 to 30 in config.py
- Add single retry with 2s backoff on ReadTimeout in kb_search tool
- Catch all exceptions and return friendly Chinese error messages instead
of propagating exceptions to the agent loop
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
- Add return_dto=None to delete_attachment decorator to fix Litestar
ImproperlyConfiguredException on 204 status code with None return type
- Catch BaseException instead of Exception in create_tables() to handle
anyio ExceptionGroup from concurrent gunicorn workers
- Enable debug=True temporarily to capture detailed error traces
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
When 2 gunicorn workers start simultaneously, both call create_all()
which races on CREATE TABLE attachments. The second worker hits a
PostgreSQL UniqueViolation on pg_type_typname_nsp_index because the
type already exists. Wrap create_tables() in try/except to handle
this gracefully.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Litestar validates that handlers with 204 status codes cannot return
a response body. The delete_attachment handler was declared as
-> Response and returned Response(content=None, status_code=204),
which caused ImproperlyConfiguredException at startup and crashed
all workers. Change return type to None and move status_code to the
@delete decorator.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Add Attachment ORM model (id, conversation_id, message_id, filename, content_type, blob_url, size_bytes, created_at) with FK to conversations
- Add POST /api/attachments/upload (multipart, 50MB limit), GET /api/attachments/{id}, GET /api/attachments/{id}/download (302 to SAS URL), DELETE /api/attachments/{id}
- Add delete_blob() and generate_sas_url() to storage/blob.py for download redirect and cleanup
- Dispatch parse_attachment task to Service Bus for parseable types (PDF, images, CSV, DOCX, XLSX)
- Pre-create blob container on startup (best-effort)
- Add AttachmentOut schema and full API docs in doc/api.md
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
Litestar's {ticket_id:uuid} path parameter requires the handler
parameter to be typed as UUID, not str. This caused a 400 validation
error on GET /api/tickets/{uuid}.
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
These packages were listed locally but never committed, causing
ModuleNotFoundError: No module named 'azure' on Azure deployment.
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
The app.storage and app.tasks packages were never committed to git,
causing ModuleNotFoundError on Azure deployment. Also adds the
document and sandbox tool modules with their config fields.
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
{ticket_id:str} was matching /api/tickets/summary before the literal route.
Changing to {ticket_id:uuid} restricts matching to UUID format only.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Replace module-level _current_model global with contextvars.ContextVar
to prevent concurrent requests from overwriting each other's search
strategy (flash vs pro)
- Add GET /api/tickets/summary returning {total, by_status, by_priority}
aggregated from Gongdan API, registered before parameterized ticket routes
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
ticketNumber (e.g. TK-2026-296691) was being used as id, causing
GET /api/tickets/{id} to 404 on upstream Gongdan API which expects UUID.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Azure Oryx could not find packages installed to .python_packages/.
Add a startup.sh script that sets PYTHONPATH to include the bundled
package directory before launching gunicorn. Update workflow to set
the Web App startup command to use startup.sh.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Add web_search tool with Jina Search/Reader/Rerank
- Flash mode: top 3, 8s timeout, no Reader/Rerank
- Pro mode: top 10, 20s timeout, concurrent Reader + Rerank top 5
- Add Redis async cache (TTL=300s) for search results
- Register "search" in ALL_TOOLS mapping
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Add complete Python backend (Litestar + LangGraph) with chat, conversations, tickets APIs
- Add GitHub Actions workflow for auto-deploying backend to Azure Web App (soc-backend)
- Add gunicorn to requirements.txt for production serving
- Update CLAUDE.md and EXTERNAL_SERVICES.md with latest config
- Remove obsolete claudehd.md (merged into gpthd.md)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>