Skip to content
Back to projects
04 2025
Applied AI Data Science

ai-memory — Segundo Cérebro (RAG)

A persistent-memory MCP server with semantic search over APIs, code, emails and schemas.

emails indexed
58k+
ERP columns catalogued
29k
searchable API endpoints
600+

From pain to result

Fig. 04 — Data flow 10 nodes · 10 edges
Emails (58k) Swagger / APIs SQL schema (29k cols) Notes & code Ingestors PostgreSQL vector + full-text Vector search Full-text (FTS) Rank fusion minimum-score cutoff MCP Server consumed by AI assistants
Source Process Store Output
01

Before

Every project started from scratch, and AI assistants guessed field names and endpoints — producing wrong code.

Every new project started from scratch: API auth flows, column names of a 29,000-field ERP, pitfalls already discovered — all scattered across emails, docs and code, and constantly forgotten.

AI assistants "guessed" field names and endpoints, producing wrong code.

02

How · 1/3

I designed a Python MCP server on PostgreSQL that ingests and indexes emails, API catalogs (Swagger), code examples, the full SQL Server schema and learning notes.

03

How · 2/3

Retrieval is a true hybrid search: it combines vector similarity (embeddings) with full-text search and fuses both rankings into one, with a minimum-score cutoff and compact output.

04

How · 3/3

It exposes MCP tools (search, detail, status, add/update note) that any compatible client — including this one — uses as the source of truth before coding.

05

After

Emails, the full ERP dictionary and hundreds of endpoints became a searchable memory, consulted before every line of code. Including while building this site.

58k+
emails indexed
29k
ERP columns catalogued
600+
searchable API endpoints

Reading the diagram

Four source families are indexed into a single database. A query splits into vector and full-text search, and both rankings meet again in a single fusion before becoming an MCP tool.

Technical highlights

  1. 1 Hybrid vector + full-text search with rank fusion and score cutoff, returning compact results referenced by [type#id].
  2. 2 Dedicated per-project ingestors rather than one generic parser: each system has its own structures (Flask routes, MongoDB schemas, API payloads), and a specific ingestor extracts what matters instead of guessing.
  3. 3 Multilingual embeddings running locally — it covers Portuguese and English with no per-query cost and without sending corporate content outside.
  4. 4 Per-project dimensional modeling, with note upsert by title to avoid duplicated knowledge.
  5. 5 Native MCP-server integration, consumable by any AI assistant — including while this portfolio was being built.