Personal Project · 2026

A warm room where AI listens, reflects, and helps you think

My Therapist is a self-hosted therapeutic companion. Not clinical software. A psychoeducational tool that combines 26 therapeutic models through an intelligent orchestrator agent, grounded in a corpus of 1,270+ documents and 54,000+ searchable knowledge chunks.

Python · FastAPI Ollama · Qdrant Vanilla HTML/CSS/JS Local-first Privacy by design

What a conversation feels like

No numbered advice lists. No forced positivity. The agent sits with the feeling first, then reflects back. Like a calm, experienced therapist in a comfortable room.

My Therapist Adaptive mode
ACT
I keep procrastinating and I hate myself for it.

There's a lot of frustration in what you're saying. Maybe some harshness toward yourself too.

Yeah. I know what I need to do, I just... can't start.

What if the inability to start isn't about willpower at all?

Sometimes when we push against ourselves that hard, there's something underneath the resistance worth looking at. What happens right before the moment you don't start?

Reflecting
What's on your mind?

How a message travels

Every message passes through four agents before reaching the user. The orchestrator routes, the RAG grounds, the therapist responds, and the safety layer reviews.

1

Orchestrator Routing

0.5 – 15s

Analyzes the message and conversation history. Outputs structured JSON: modality, themes, emotional tone, intent, language. 26 possible routes. No prose, no guessing — deterministic and testable.

2

RAG Retrieval

0.3 – 1s

Embeds the query, searches Qdrant filtered by the chosen modality. Top 4 chunks above 0.3 similarity. If nothing matches the modality, falls back to unfiltered search. The therapist is grounded in real psychoeducational material, not just its training data.

3

Therapist Response

1 – 50s

Loads the modality-specific system prompt (26 hand-crafted prompts teaching how real practitioners in that school actually talk). Injects RAG context and conversation history. Generates a response streamed via SSE with heartbeats every 8 seconds.

4

Safety Review (3 tiers)

0 – 2s

Tier 1: Deterministic regex guardrails scan every user message for crisis signals (EN/FR/NL) and every AI response for diagnostic language, prescriptive advice, identity claims, and hallucination patterns. Zero latency, no model dependency. Tier 2: A separate LLM agent reviews the response for subtler safety issues. Tier 3: Universal safety directives are injected into all 26 sub-agent prompts. Defense in depth — three independent layers.

Crisis Protocol

If the safety layer detects suicidal ideation or acute distress, it immediately replaces the AI response with verified emergency contacts localized to the user's language. No diagnosis, no treatment, no delay. Every agent prompt explicitly guards: "Never diagnose. Never prescribe. Never claim to be a therapist."

Emergency Resources Worldwide

Belgium 0800 32 123 Centre de Prévention du Suicide
France 3114 Numéro National de Prévention du Suicide
United States 988 Suicide & Crisis Lifeline
United Kingdom 116 123 Samaritans (24/7)
Canada 988 Suicide Crisis Helpline
Netherlands 113 Zelfmoordpreventie
Germany 0800 111 0 111 Telefonseelsorge
US, Canada, UK & Ireland Text HOME to 741741 Crisis Text Line (24/7, free)

Full list of 175+ countries: findahelpline.com · IASP Crisis Centres

Built-in Guardrails

  • Tier 1 — Regex: Deterministic pattern matching catches crisis signals, diagnostic language, and hallucinated citations before any LLM runs
  • Tier 2 — LLM Safety Agent: Separate model call reviews every response (temperature 0.1 for precision)
  • Tier 3 — Prompt Directives: Anti-hallucination and boundary rules injected into all 26 sub-agent prompts
  • Red flags (diagnosis, prescription, identity claims) trigger full response replacement
  • Yellow flags (medication mentions, fabricated statistics) trigger disclaimers
  • All conversations are local-only — SQLite on-device, never transmitted externally

26 specialized modalities

Each modality has a hand-crafted system prompt. Not generic roleplay — authentic representations of how practitioners in each school actually work. The orchestrator picks the best fit for each message.

Cognitive

  • CBT
  • ACT
  • DBT
  • Schema Therapy
  • Brief Therapy

Depth Work

  • IFS (Internal Family Systems)
  • Psychoanalysis
  • Gestalt
  • Narrative Therapy
  • Attachment Theory
  • EMDR Concepts

Body & Emotions

  • Somatic Approaches
  • Positive Psychology
  • Motivational Interviewing
  • Resilience

Communication

  • NVC (Nonviolent Communication)
  • Transactional Analysis
  • PCM (Process Communication)
  • Systemic Therapy

Growth & Meaning

  • Enneagram
  • NLP
  • Spiral Dynamics
  • Growth Mindset
  • Logotherapy (Frankl)
  • Stoicism
  • Leadership Psychology

Scale of the system

26
Modalities
1,270+
Source Documents
54k+
RAG Chunks
13
Cloud Fallbacks
8
Dependencies
4
Agent Pipeline
Fallback chain: Cerebras Groq SambaNova OpenRouter Ollama (local)

What makes this interesting

Modality-Aware RAG

The vector search isn't just "find similar text." It filters by the therapeutic modality the orchestrator selected. CBT agent pulls CBT corpus. Schema Therapy pulls Schema corpus. Context stays relevant and aligned.

SSE with Heartbeats

Cloud proxies kill idle connections after ~25 seconds. The backend sends heartbeat events every 8 seconds to keep the stream alive during slow inference. The user sees "Reflecting..." instead of a timeout error.

Safety as Separate Agent

Rather than baking safety rules into the therapist prompt (where they can be overridden), a completely separate LLM call reviews every response. Defense in depth. The safety agent can veto or rewrite anything.

Smart Cloud Fallback

13 free-tier providers with automatic per-request failover. One provider returns a 429? Next one is tried transparently. The user experiences slightly slower responses but never sees an error message.

Resume-Safe Pipeline

The JSONL checkpoint means you can kill and restart the embedding process at any time without losing work. Merge-on-save, not overwrite. At 0.6 chunks/second on CPU, that matters for 54,000+ chunks.

Minimal Dependencies

8 Python packages. No LangChain, no LlamaIndex, no heavy frameworks. Direct HTTP calls to Ollama and Qdrant. Full control, full transparency, easy to debug. The entire backend is ~2,500 lines of focused Python.

Ilse Crawford meets Thich Nhat Hanh

The entire visual language is rooted in specific designers and thinkers. Warm linen, sage green, soft gold. The interface breathes. No purple gradients, no gamification, no clinical feel.

Design Principles

  • Warm, organic, unhurried
  • Messages fade in gently (800ms)
  • Brown-tinted shadows, never pure black
  • Lora serif for conversations, generous line-height
  • Typing indicator breathes slowly (2.4s cycle)
  • Japanese "Ma" spacing — meaningful emptiness
  • Respects prefers-reduced-motion

Deliberately Avoided

  • Purple gradients and neon accents
  • Gamification, streaks, and badges
  • Clinical forms and questionnaires
  • "Daily check-in" notifications
  • Progress tracking dashboards
  • Loading spinners (use "Reflecting..." instead)
  • Avatars and character illustrations

Technology choices

Local-first architecture. Everything runs on a single Linux machine. Cloud is an escape hatch for speed, not a dependency.

Python
Backend language
FastAPI
API framework
Ollama
Local LLM inference
Qdrant
Vector database
SQLite
Conversation storage
nomic-embed-text
Embeddings model
Vanilla JS
Frontend (no framework)
SSE
Real-time streaming

Want to experience it yourself?

My Therapist runs on a private server. Access is gated through Cloudflare to keep it personal and secure. Leave your email below and I'll send you an invite within 24 hours.

Your email is only used to grant access. No spam, no mailing list.

"We always prioritize the experience over the way something looks."
Ilse Crawford