Dolphin is a memory layer that combines vector search with a knowledge graph. Use it as a Python SDK inside your LLM application, or as shared memory for Claude Code and other local agents, where what one session learns the next one already knows.
| Memory for your LLM app | Shared memory for agents | |
|---|---|---|
| What | A Python SDK: add, search, get_context |
An MCP server and a Claude Code plugin |
| Who remembers | Your app's users, one namespace per user_id |
Every agent session working on the same project |
| Storage | Local SQLite file, or Supabase for server-side apps | Local SQLite file shared by all sessions on the machine |
| Start here | SDK quick start | Agent quick start |
Full reference: sdk/README.md. Internals: ARCHITECTURE.md.
Note
Version 0.2.0 (local SQLite storage, MCP server, Claude Code plugin) is not on PyPI yet; pip install dolphin-memory currently gives 0.1.0, which requires Supabase. Until it is published, install from source:
pip install "git+https://github.com/DewashishCodes/dolphin@main#subdirectory=sdk"pip install dolphin-memoryfrom dolphin_memory import DolphinMemory
# Local by default: memories live in ~/.dolphin/dolphin.db. No account, no server.
memory = DolphinMemory()
memory.add("I'm a software engineer in Mumbai. I love rock climbing.", user_id="u1")
context = memory.get_context("Suggest a weekend activity", user_id="u1")
# ### RELEVANT MEMORIES
# - [Just now]: I'm a software engineer in Mumbai. I love rock climbing.
#
# ### KNOWLEDGE GRAPH
# User LIKES Rock Climbing (Sport)
# User LIVES_IN Mumbai (City)Inject context into your system prompt. For server-side apps, install dolphin-memory[supabase] and pass supabase_url / supabase_key; the API is the same.
Storing and searching memories works out of the box. Building the knowledge graph needs an LLM, by default a local Ollama model:
dolphin setup # installs Ollama and pulls llama3.2
dolphin doctor # checks that everything is healthypip install dolphin-memory/plugin marketplace add DewashishCodes/dolphin
/plugin install dolphin-memory@dolphin
From then on, in every project:
| When | What happens |
|---|---|
| A session starts | A short brief of the project's most recent memories is added to context |
| You send a prompt | Memories related to that prompt are recalled and added to it |
| Claude finishes responding | The exchange is distilled; decisions, conventions and gotchas are stored, routine work is not |
Claude also gets remember, recall, search, forget and graph_stats tools, and the /dolphin-memory:recall, /dolphin-memory:remember and /dolphin-memory:status commands.
Memories are scoped to the project (its git remote), so every clone and every session of one repository shares them. Open a second terminal, start a new session, and it knows what the first one decided.
Only want the tools, or using another MCP client? Skip the plugin:
claude mcp add dolphin -- dolphin mcp- Semantic layer. Every memory is embedded and stored. Near-duplicates reinforce the existing memory instead of piling up.
- Graph layer. An LLM extracts
(subject) -[RELATIONSHIP]-> (object)triples in the background and merges them into a knowledge graph. - Hybrid recall. A query finds similar memories, then similar graph nodes, then walks one hop out from those nodes, so related facts surface even when they share no words with the query.
- Consolidation.
consolidate()asks the LLM to merge duplicate graph nodes ("Bill Gates", "William Gates III").
Everything runs locally by default: SQLite for storage, an ONNX embedding model, and Ollama for extraction. Gemini and OpenAI are optional extraction providers.
| Path | What it is |
|---|---|
sdk/ |
The dolphin-memory Python package: SDK, MCP server, daemon, hooks, CLI |
plugin/ |
The Claude Code plugin (hooks, MCP config, skills) |
.claude-plugin/ |
Marketplace manifest that publishes plugin/ |
docs/ |
Landing page (Firebase Hosting) |
server.py, database/, static/ |
The original Dolphin chat demo with a 3D graph view; see below |
python -m venv venv
venv/Scripts/python -m pip install -e "sdk[dev]" # bin/ instead of Scripts/ on macOS/Linux
cd sdk && ../venv/Scripts/python -m pytest -qThe tests use a fake embedder and a stubbed extractor, so they need no model downloads, Ollama, or Supabase.
Dolphin started as a chat app that builds a global knowledge graph and renders it as an interactive 3D force graph. It still lives at the repository root and runs on Supabase + Gemini/Ollama, separately from the SDK.
dolphinDemo.1.mp4
Run the chat demo
- Create a Supabase project and run
supabase/migrations/20260214062656_remote_schema.sql, thensupabase/migrations/20260412000000_v2_critical_fixes.sql, in the SQL Editor. - Copy
.env.exampleto.envand fill inSUPABASE_URL,SUPABASE_KEYandGOOGLE_API_KEY. - Install Ollama and run
ollama pull llama3.2. pip install -r requirements.txtuvicorn server:app --reload, then open http://localhost:8000.
-
dolphin-memorySDK on PyPI - Local SQLite backend, no setup
- MCP server
- Claude Code plugin with automatic recall and capture
- Claude as an extraction provider
- Capture at session end and before compaction
- Recall ranking by recency and reinforcement
- Graph viewer for the local store
- Dolphin Cloud: managed, team-shared memory
Made with ❤️ by DewashishCodes
