RFC: move the chatbot vector store from Zilliz to Neon
The AI bubble chat, called the chatbot here, keeps its knowledge in a free Zilliz cluster that stops when nobody uses it. This RFC moves that knowledge to Postgres with pgvector on Neon, so all our databases live with one provider.
Recommendation: a new chatbot project in the Skill Pixel Neon organization (option A). Beta and production share one set of chunks, filtered by tenant, as they do today.
One design decision is open, plus three questions about owners and the Neon plan. They are listed under Open questions.
Why
- The production Zilliz cluster is on the free tier and stops after a quiet period. We found it stopped on 2026-10-09.
- While it is stopped, every knowledge search returns nothing and every document upload fails. Someone has to resume it by hand.
- On 2026-10-10 Steven asked on SP-593 for a more stable database and suggested Neon.
- The LMS already runs on Neon. Keeping the chatbot's data there too gives us one provider to operate, monitor and pay.
Alternatives considered
| Alternative | Why not |
|---|---|
| Paid Zilliz tier | Fixes the stopping, but keeps a second database provider next to Neon. |
| Keep-alive ping to Zilliz | Works around the free tier instead of fixing it, and still leaves two providers. |
Goals
- Knowledge search keeps working after idle periods with no manual step. Cold start latency is a known gap, see Risks.
- The chatbot's data sits with the rest of our databases on Neon.
- We can see what the vector store uses and get warned before a limit is hit.
- Every query stays limited to the calling tenant.
Not in this RFC
- Separating beta and production knowledge. Both LMS environments keep sharing one set of chunks, filtered by tenant.
- Separating beta and production chat history. MongoDB stays shared.
- Limiting the HistoryClass tutor to one course (SP-714). That gets its own ticket.
- Copying vectors out of Zilliz. It holds one document, which we upload again.
Current state
Chatbot
- The chatbot (repo
lms-ragflow) is one Railway service. Beta and production LMS both call it. - Zilliz holds one tenant collection with one document, the SkillPixel knowledge base PDF, embedded with
text-embedding-3-largeat 3072 dimensions. - Requests without a valid
X-Tenant-IDalready get400. Callers cannot pass their own filter.
Neon account
Read from the Neon API on 2026-10-11.
| Item | Value |
|---|---|
| Organization | Skill Pixel, plan free_v3 |
| Project | LMS platform (lively-fire-58138144), ap-southeast-1, Postgres 17 |
| Branches | production only. Beta LMS runs on Supabase. |
| Storage | 283 MB of the 1 GB free limit |
| Usage since 2026-10-10, about 34 hours | 6.2 CU-hours (compute unit hours), 303 MB of egress |
| pgvector | 0.8.0 available, not installed |
The LMS itself may hit its free limits this month. SP-744 says the organization moved to the Launch plan, but the API reports the free plan. At the current pace the LMS would use about 135 CU-hours and 6.7 GB of egress this month. The free plan allows 100 CU-hours and 5 GB per project, and suspends the compute until next month when either runs out.
This is an estimate from 34 hours of data. It affects the LMS whether or not this RFC goes ahead, so it is listed as its own question below.
Neon plans, the parts that matter here
| Free | Launch | |
|---|---|---|
| Compute | 100 CU-hours per project | $0.106 per CU-hour |
| Storage | 1 GB per project. Writes are blocked above it. | $0.35 per GB-month, no limit |
| Egress | 5 GB per project | 500 GB per project, then $0.10 per GB |
| Scale to zero | After 5 idle minutes. Cannot be turned off. | After 5 idle minutes. Can be turned off. |
| Projects | 100 per organization | 100 per organization |
| Cost alerts | None | Spending notifications and limits, see Monitoring |
Source: neon.com/pricing
Design
This part is the same whichever project we pick.
- One table, filtered by tenant. All tenants share
document_chunks. Every statement the store runs filters on the tenant fromX-Tenant-ID, and a test checks that each statement binds it. Beta and production read the same rows. - Embeddings. Stored as
halfvec(3072)with an HNSW cosine index, because pgvector indexes plain vectors only up to 2000 dimensions. I am checking on the real PDF whethervector(1536)from the same model retrieves as well at half the size. The result will replace this line. - Least-privilege role. The chatbot connects as
chatbot_rw, which can only select, insert and delete rows. The database owner runs the schema script. The table has no sequence, so the missing sequence grant from the LMS's Neon incident cannot happen here. - Safe replace. Replacing a document is one transaction, locked per tenant and document, so a failed or parallel upload never leaves the document empty or mixed. With Zilliz, a failed upload left it empty.
Connections and search
| Setting | Value | Why |
|---|---|---|
| Endpoint | Neon direct host, not -pooler | The pooler runs PgBouncer in transaction mode, which breaks asyncpg's prepared statements. |
| Pool | asyncpg, 0 to 5 connections, idle ones closed after 60 s | Small enough for a 0.25 CU compute, and idle connections do not keep it awake. |
| Connect timeout | 10 s per attempt | A waking compute answers well within this in normal cases. |
| Retried errors | TooManyConnectionsError, CannotConnectNowError, ConnectionDoesNotExistError, OSError, connect timeouts | The errors Neon's proxy returns while a compute wakes up (SP-625). Other errors are not retried. |
| Retry budget | 3 attempts, waits of 0.5 s then 1 s, 25 s total | Stays under the LMS's 30 s chat timeout. |
| Query timeout | 10 s per statement | A stuck query cannot hold a connection. |
| When retries run out | The search tool tells the agent the knowledge store is unavailable, and logs an error | An empty result would look like "no knowledge", the same silent failure as stopped Zilliz. |
| Search settings | SET LOCAL hnsw.iterative_scan = relaxed_order and hnsw.ef_search = 100 inside each search transaction | The index searches all tenants first, then filters. Iterative scans keep reading until a small tenant has its full result count. |
| Columns returned | id, document, text, metadata and score, never the embedding | The embedding is about 6 KB per chunk and nobody reads it back. |
Decision: which Neon project
| A. New chatbot project (recommended) | B. Database inside LMS platform | |
|---|---|---|
| Limits | Its own 1 GB, 100 CU-hours and 5 GB of egress | Shares the LMS's limits, which the LMS is already on pace to exceed |
| If a limit is hit | Only the chatbot stops. An LMS limit does not stop the chatbot either. | The LMS and the chatbot stop together. |
| Cold starts | More often. The compute scales to zero after 5 idle minutes. | Fewer. Learner traffic keeps the compute awake. |
| Cost visibility | Chatbot usage is its own number. | Mixed with LMS usage. Neon has no per-database split. |
| One place for databases | Same Neon organization, separate project | Same Neon project |
Recommendation: A. Both keep every database on Neon. A also stops the chatbot and the LMS from taking each other down, at no extra cost. If the organization moves to Launch later, the chatbot stays on A and can turn off scale to zero.
Cost and monitoring
Expected chatbot usage
These are estimates. There is one compute, because beta and production share it.
| Estimate | Free plan room | |
|---|---|---|
| Storage | The PDF is 63 chunks. At about 14 KB each with text and index, that is under 1 MB. | 1 GB, including history kept for restores |
| Compute | Each wake-up runs at least 5 minutes at 0.25 CU. About 20 separate chat sessions a day, each keeping the compute up 10 minutes, is about 25 CU-hours a month. Always on would be about 180. | 100 CU-hours |
| Egress | About 10 KB per search: five chunks of text, no embeddings | About 500,000 searches in 5 GB |
| On Launch | About $3 a month at that pattern, about $19 if always on |
OpenAI embedding cost for uploads already reaches the LMS usage ledger (SP-797). Uploading the one document again costs a few cents.
How we watch it
- Neon Console. The Projects page and each project's Overview show compute, storage, history and network transfer for the month, up to an hour behind. The free plan keeps monitoring data for one day.
- Usage check. Neon sends no alerts on the free plan. A scheduled GitHub Actions workflow in
lms-ragflowruns every hour, reads the project's usage from the Neon API (compute_time_seconds,data_transfer_bytes,synthetic_storage_size) and posts to Discord at 80% of a limit. It also posts when the check itself fails. Owner: Khai Hoan. - Health check. The same workflow runs one known search for SkillPixel and alerts when it fails or returns nothing.
- On Launch. Spending notifications email organization admins at 80% and 100% of an amount they set, without stopping usage. Per-project consumption limits can suspend compute at a cap.
Risks and notes
| Risk | Plan |
|---|---|
| A cold start takes longer than the LMS's 30 s chat timeout. SP-625 saw 33 to 38 s on the LMS project. | Noted, solved later. The connection retry covers short wake-ups only. Options for later: measure the chatbot's own cold start, a warm-up call when the chat bubble opens, a longer LMS timeout, or turning off scale to zero on Launch. |
| Beta and production share chunks, so a beta upload changes what production learners see. | Accepted. Upload through production only, or agree before uploading from beta. |
| A free plan limit suspends the chatbot until next month. | Own project (option A) and the hourly usage check. |
| The runtime role lacks a privilege on a new object. | The schema script grants explicitly, and the chatbot never runs DDL. |
| One missed tenant filter in a future query would expose every tenant. | All queries go through one store class, with a test that each statement binds the tenant. Row-level security is an open question. |
Rollout
Assuming option A.
- A Neon admin creates the chatbot project, the
chatbotdatabase, and the owner andchatbot_rwroles, then runs the schema script. - Ship the chatbot change on SP-811 and set
VECTOR_DATABASE_URLon Railway. - Upload the SkillPixel knowledge base PDF right after the deploy, outside learner hours.
- Ask a question the PDF answers, and check that the reply cites its chunks.
- Take a Neon snapshot, then remove the Zilliz variables and delete the cluster.
Open questions
- Open Which Neon project: option A, a new chatbot project, or option B, a database inside LMS platform.
- Open Should the organization move to Launch for the LMS's own compute and egress limits, and who confirms the plan SP-744 describes?
- Open Who creates the Neon project and roles? My access to LMS platform is editor, not admin.
- Open Do we want row-level security on
document_chunksas a second guard behind the tenant filter? - Decided Beta and production share one set of chunks, filtered by tenant.
- Decided No migration or rollback plan. Zilliz holds one document, which we upload again.