RFC: move the chatbot vector store from Zilliz to Neon
The AI bubble chat, called the chatbot here, keeps its knowledge in a free Zilliz cluster that stops when nobody uses it. This RFC moves that knowledge to Postgres with pgvector on Neon, so all our databases live with one provider.
Recommendation: a new chatbot project in the Skill Pixel Neon organization (option A). Beta and production share one set of chunks, filtered by tenant, as they do today.
One design decision is open, plus one question about the Neon plan. They are listed under Open questions.
Why
- The production Zilliz cluster is on the free tier and stops after a quiet period. We found it stopped on 2026-10-09.
- While it is stopped, every knowledge search returns nothing and every document upload fails. Someone has to resume it by hand.
- On 2026-10-10 Steven asked on SP-593 for a more stable database and suggested Neon.
- The LMS already runs on Neon. Keeping the chatbot's data there too gives us one provider to operate, monitor and pay.
Alternatives considered
| Alternative | Why not |
|---|---|
| Paid Zilliz tier | Fixes the stopping, but keeps a second database provider next to Neon. |
| Keep-alive ping to Zilliz | Works around the free tier instead of fixing it, and still leaves two providers. |
Goals
- Knowledge search keeps working after idle periods with no manual step. Cold start latency is a known gap, see Risks.
- The chatbot's data sits with the rest of our databases on Neon.
- We can see what the vector store uses and get warned before a limit is hit.
- Every query stays limited to the calling tenant.
Not in this RFC
- Separating beta and production knowledge. Both LMS environments keep sharing one set of chunks, filtered by tenant.
- Separating beta and production chat history. MongoDB stays shared.
- Limiting the HistoryClass tutor to one course (SP-714). That gets its own ticket.
- Copying vectors out of Zilliz. It holds one document, which we upload again.
Current state
Chatbot
- The chatbot (repo
lms-ragflow) is one Railway service. Beta and production LMS both call it. - Zilliz holds one tenant collection with one document, the SkillPixel knowledge base PDF, embedded with
text-embedding-3-largeat 3072 dimensions. - Requests without a valid
X-Tenant-IDalready get400. Callers cannot pass their own filter.
Neon account
Read from the Neon API on 2026-10-11, checked twice.
| Item | Value |
|---|---|
| Organization | Skill Pixel, plan free (free_v3 on the project record) |
| Project | LMS platform (lively-fire-58138144), ap-southeast-1, Postgres 17 |
| Branches | production only. Beta LMS runs on Supabase. |
| Storage | 283 MB of the 1 GB free limit |
| Billing period | 2026-10-10 03:16 UTC to 2026-11-01 00:00 UTC, about 22 days |
| Usage in the first 36 hours | 6.65 CU-hours (compute unit hours), 323 MB of egress |
| pgvector | 0.8.0 available, not installed |
The organization is on the free plan, not Launch. SP-744 (created 2026-09-23) says Neon was upgraded to Launch with a card on file. On 2026-10-11 the Neon API reports the plan as free, both on the organization and on the LMS project. The current billing period started on 2026-10-10 at 03:16 UTC, the same moment the organization record was last changed, which suggests the plan changed that day. The API does not say why.
The LMS is close to its free limits. At the pace of the first 36 hours, the LMS uses about 97 of 100 CU-hours and 4.7 of 5 GB of egress by 2026-11-01. A full 30-day month at that pace is about 132 CU-hours and 6.4 GB, over both. When either runs out on the free plan, the compute is suspended until the next period, and the LMS goes down.
This affects the LMS whether or not this RFC goes ahead, so it is listed as its own question below. Only an organization admin can see or change billing.
Neon plans, the parts that matter here
| Free | Launch | |
|---|---|---|
| Compute | 100 CU-hours per project | $0.106 per CU-hour |
| Storage | 1 GB per project. Writes are blocked above it. | $0.35 per GB-month, no limit |
| Egress | 5 GB per project | 500 GB per project, then $0.10 per GB |
| Scale to zero | After 5 idle minutes. Cannot be turned off. | After 5 idle minutes. Can be turned off. |
| Projects | 100 per organization | 100 per organization |
| Cost alerts | None | Spending notifications and limits, see Monitoring |
Source: neon.com/pricing
Design
This part is the same whichever project we pick.
- One table, filtered by tenant. All tenants share
document_chunks. Every statement the store runs filters on the tenant fromX-Tenant-ID, and a test checks that each statement binds it. Beta and production read the same rows. - Embeddings. Stored as
halfvec(3072)with an HNSW cosine index, because pgvector indexes plain vectors only up to 2000 dimensions. I am checking on the real PDF whethervector(1536)from the same model retrieves as well at half the size. The result will replace this line. - Least-privilege role. The chatbot connects as
chatbot_rw, which can only select, insert and delete rows. The database owner runs the schema script. The table has no sequence, so the missing sequence grant from the LMS's Neon incident cannot happen here. - Safe replace. Replacing a document is one transaction, locked per tenant and document, so a failed or parallel upload never leaves the document empty or mixed. With Zilliz, a failed upload left it empty.
Connections and search
| Setting | Value | Why |
|---|---|---|
| Endpoint | Neon direct host, not -pooler | The pooler runs PgBouncer in transaction mode, which breaks asyncpg's prepared statements. |
| Pool | asyncpg, 0 to 5 connections, idle ones closed after 60 s | Small enough for a 0.25 CU compute, and idle connections do not keep it awake. |
| Connect timeout | 10 s per attempt | A waking compute answers well within this in normal cases. |
| Retried errors | TooManyConnectionsError, CannotConnectNowError, ConnectionDoesNotExistError, OSError, connect timeouts | The errors Neon's proxy returns while a compute wakes up (SP-625). Other errors are not retried. |
| Retry budget | 3 attempts, waits of 0.5 s then 1 s, 25 s total | Stays under the LMS's 30 s chat timeout. |
| Query timeout | 10 s per statement | A stuck query cannot hold a connection. |
| When retries run out | The search tool tells the agent the knowledge store is unavailable, and logs an error | An empty result would look like "no knowledge", the same silent failure as stopped Zilliz. |
| Search settings | SET LOCAL hnsw.iterative_scan = relaxed_order and hnsw.ef_search = 100 inside each search transaction | The index searches all tenants first, then filters. Iterative scans keep reading until a small tenant has its full result count. |
| Columns returned | id, document, text, metadata and score, never the embedding | The embedding is about 6 KB per chunk and nobody reads it back. |
Decision: which Neon project
| A. New chatbot project (recommended) | B. Database inside LMS platform | |
|---|---|---|
| Limits | Its own 1 GB, 100 CU-hours and 5 GB of egress | Shares the LMS's limits, which the LMS is already on pace to exceed |
| If a limit is hit | Only the chatbot stops. An LMS limit does not stop the chatbot either. | The LMS and the chatbot stop together. |
| Cold starts | More often. The compute scales to zero after 5 idle minutes. | Fewer. Learner traffic keeps the compute awake. |
| Cost visibility | Chatbot usage is its own number. | Mixed with LMS usage. Neon has no per-database split. |
| One place for databases | Same Neon organization, separate project | Same Neon project |
Recommendation: A. Both keep every database on Neon. A also stops the chatbot and the LMS from taking each other down, at no extra cost. If the organization moves to Launch later, the chatbot stays on A and can turn off scale to zero.
Cost and monitoring
Expected chatbot usage
These are estimates. There is one compute, because beta and production share it.
| Estimate | Free plan room | |
|---|---|---|
| Storage | The PDF is 63 chunks. At about 14 KB each with text and index, that is under 1 MB. | 1 GB, including history kept for restores |
| Compute | Each wake-up runs at least 5 minutes at 0.25 CU. About 20 separate chat sessions a day, each keeping the compute up 10 minutes, is about 25 CU-hours a month. Always on would be about 180. | 100 CU-hours |
| Egress | About 10 KB per search: five chunks of text, no embeddings | About 500,000 searches in 5 GB |
| On Launch | About $3 a month at that pattern, about $19 if always on |
OpenAI embedding cost for uploads already reaches the LMS usage ledger (SP-797). Uploading the one document again costs a few cents.
How we watch it
- Neon Console. The Projects page and each project's Overview show compute, storage, history and network transfer for the month, up to an hour behind. The free plan keeps monitoring data for one day.
- Usage check. Neon sends no alerts on the free plan. A scheduled GitHub Actions workflow in
lms-ragflowruns every hour, reads the project's usage from the Neon API (compute_time_seconds,data_transfer_bytes,synthetic_storage_size) and posts to Discord at 80% of a limit. It also posts when the check itself fails. Owner: Khai Hoan. - Health check. The same workflow runs one known search for SkillPixel and alerts when it fails or returns nothing.
- On Launch. Spending notifications email organization admins at 80% and 100% of an amount they set, without stopping usage. Per-project consumption limits can suspend compute at a cap.
Risks and notes
| Risk | Plan |
|---|---|
| A cold start takes longer than the LMS's 30 s chat timeout. SP-625 saw 33 to 38 s on the LMS project. | Noted, solved later. The connection retry covers short wake-ups only. Options for later: measure the chatbot's own cold start, a warm-up call when the chat bubble opens, a longer LMS timeout, or turning off scale to zero on Launch. |
| Beta and production share chunks, so a beta upload changes what production learners see. | Accepted. Upload through production only, or agree before uploading from beta. |
| A free plan limit suspends the chatbot until next month. | Own project (option A) and the hourly usage check. |
| The runtime role lacks a privilege on a new object. | The schema script grants explicitly, and the chatbot never runs DDL. |
| One missed tenant filter in a future query would expose every tenant. | All queries go through one store class, with a test that each statement binds the tenant. |
Rollout
Assuming option A.
- An organization admin or editor creates the chatbot project and grants Khai Hoan editor on it. Khai Hoan then creates the
chatbotdatabase, the owner andchatbot_rwroles, and runs the schema script. - Ship the chatbot change on SP-811 and set
VECTOR_DATABASE_URLon Railway. - Upload the SkillPixel knowledge base PDF right after the deploy, outside learner hours.
- Ask a question the PDF answers, and check that the reply cites its chunks.
- Take a Neon snapshot, then remove the Zilliz variables and delete the cluster.
Open questions
- Open Which Neon project: option A, a new chatbot project, or option B, a database inside LMS platform.
- Open The organization admin (SkillPixel Admin account) confirms whether the SP-744 upgrade to Launch is still in place, and whether to move to Launch before the LMS reaches its free limits.
- Answered Who creates the Neon project and roles. Khai Hoan is a collaborator in the Skill Pixel organization, and Neon does not let collaborators create projects. For option A, SkillPixel Admin (organization admin) or Thắng Phạm (organization editor) creates the project and grants Khai Hoan editor on it. Khai Hoan then creates the database, roles and schema. For option B, Khai Hoan is already editor on LMS platform, which allows creating databases and Postgres roles and running SQL.
- Decided Beta and production share one set of chunks, filtered by tenant.
- Decided No migration or rollback plan. Zilliz holds one document, which we upload again.