RFC: move the chatbot vector store from Zilliz to Neon

Draft for review Updated 2026-10-11 Author: Khai Hoan Reviewers: Khang Nguyen, Steven Ngo Jira: SP-811

The AI bubble chat, called the chatbot here, keeps its knowledge in a free Zilliz cluster that stops when nobody uses it. This RFC moves that knowledge to Postgres with pgvector on Neon, so all our databases live with one provider.

Recommendation: a new chatbot project in the Skill Pixel Neon organization (option A). Beta and production share one set of chunks, filtered by tenant, as they do today.

One design decision is open, plus three questions about owners and the Neon plan. They are listed under Open questions.

Why

Alternatives considered

AlternativeWhy not
Paid Zilliz tierFixes the stopping, but keeps a second database provider next to Neon.
Keep-alive ping to ZillizWorks around the free tier instead of fixing it, and still leaves two providers.

Goals

Not in this RFC

Current state

Chatbot

Neon account

Read from the Neon API on 2026-10-11.

ItemValue
OrganizationSkill Pixel, plan free_v3
ProjectLMS platform (lively-fire-58138144), ap-southeast-1, Postgres 17
Branchesproduction only. Beta LMS runs on Supabase.
Storage283 MB of the 1 GB free limit
Usage since 2026-10-10, about 34 hours6.2 CU-hours (compute unit hours), 303 MB of egress
pgvector0.8.0 available, not installed

The LMS itself may hit its free limits this month. SP-744 says the organization moved to the Launch plan, but the API reports the free plan. At the current pace the LMS would use about 135 CU-hours and 6.7 GB of egress this month. The free plan allows 100 CU-hours and 5 GB per project, and suspends the compute until next month when either runs out.

This is an estimate from 34 hours of data. It affects the LMS whether or not this RFC goes ahead, so it is listed as its own question below.

Neon plans, the parts that matter here

FreeLaunch
Compute100 CU-hours per project$0.106 per CU-hour
Storage1 GB per project. Writes are blocked above it.$0.35 per GB-month, no limit
Egress5 GB per project500 GB per project, then $0.10 per GB
Scale to zeroAfter 5 idle minutes. Cannot be turned off.After 5 idle minutes. Can be turned off.
Projects100 per organization100 per organization
Cost alertsNoneSpending notifications and limits, see Monitoring

Source: neon.com/pricing

Design

This part is the same whichever project we pick.

Connections and search

SettingValueWhy
EndpointNeon direct host, not -poolerThe pooler runs PgBouncer in transaction mode, which breaks asyncpg's prepared statements.
Poolasyncpg, 0 to 5 connections, idle ones closed after 60 sSmall enough for a 0.25 CU compute, and idle connections do not keep it awake.
Connect timeout10 s per attemptA waking compute answers well within this in normal cases.
Retried errorsTooManyConnectionsError, CannotConnectNowError, ConnectionDoesNotExistError, OSError, connect timeoutsThe errors Neon's proxy returns while a compute wakes up (SP-625). Other errors are not retried.
Retry budget3 attempts, waits of 0.5 s then 1 s, 25 s totalStays under the LMS's 30 s chat timeout.
Query timeout10 s per statementA stuck query cannot hold a connection.
When retries run outThe search tool tells the agent the knowledge store is unavailable, and logs an errorAn empty result would look like "no knowledge", the same silent failure as stopped Zilliz.
Search settingsSET LOCAL hnsw.iterative_scan = relaxed_order and hnsw.ef_search = 100 inside each search transactionThe index searches all tenants first, then filters. Iterative scans keep reading until a small tenant has its full result count.
Columns returnedid, document, text, metadata and score, never the embeddingThe embedding is about 6 KB per chunk and nobody reads it back.

Decision: which Neon project

A. New chatbot project (recommended)B. Database inside LMS platform
LimitsIts own 1 GB, 100 CU-hours and 5 GB of egressShares the LMS's limits, which the LMS is already on pace to exceed
If a limit is hitOnly the chatbot stops. An LMS limit does not stop the chatbot either.The LMS and the chatbot stop together.
Cold startsMore often. The compute scales to zero after 5 idle minutes.Fewer. Learner traffic keeps the compute awake.
Cost visibilityChatbot usage is its own number.Mixed with LMS usage. Neon has no per-database split.
One place for databasesSame Neon organization, separate projectSame Neon project

Recommendation: A. Both keep every database on Neon. A also stops the chatbot and the LMS from taking each other down, at no extra cost. If the organization moves to Launch later, the chatbot stays on A and can turn off scale to zero.

Cost and monitoring

Expected chatbot usage

These are estimates. There is one compute, because beta and production share it.

EstimateFree plan room
StorageThe PDF is 63 chunks. At about 14 KB each with text and index, that is under 1 MB.1 GB, including history kept for restores
ComputeEach wake-up runs at least 5 minutes at 0.25 CU. About 20 separate chat sessions a day, each keeping the compute up 10 minutes, is about 25 CU-hours a month. Always on would be about 180.100 CU-hours
EgressAbout 10 KB per search: five chunks of text, no embeddingsAbout 500,000 searches in 5 GB
On LaunchAbout $3 a month at that pattern, about $19 if always on

OpenAI embedding cost for uploads already reaches the LMS usage ledger (SP-797). Uploading the one document again costs a few cents.

How we watch it

Risks and notes

RiskPlan
A cold start takes longer than the LMS's 30 s chat timeout. SP-625 saw 33 to 38 s on the LMS project.Noted, solved later. The connection retry covers short wake-ups only. Options for later: measure the chatbot's own cold start, a warm-up call when the chat bubble opens, a longer LMS timeout, or turning off scale to zero on Launch.
Beta and production share chunks, so a beta upload changes what production learners see.Accepted. Upload through production only, or agree before uploading from beta.
A free plan limit suspends the chatbot until next month.Own project (option A) and the hourly usage check.
The runtime role lacks a privilege on a new object.The schema script grants explicitly, and the chatbot never runs DDL.
One missed tenant filter in a future query would expose every tenant.All queries go through one store class, with a test that each statement binds the tenant. Row-level security is an open question.

Rollout

Assuming option A.

  1. A Neon admin creates the chatbot project, the chatbot database, and the owner and chatbot_rw roles, then runs the schema script.
  2. Ship the chatbot change on SP-811 and set VECTOR_DATABASE_URL on Railway.
  3. Upload the SkillPixel knowledge base PDF right after the deploy, outside learner hours.
  4. Ask a question the PDF answers, and check that the reply cites its chunks.
  5. Take a Neon snapshot, then remove the Zilliz variables and delete the cluster.

Open questions