When I built NaLog Agent (a Qwen-powered memory-augmented agent for smallholder rice farmers), the hardest part was memory management. A farmer asks “Does Paddy 3 need pumping now?” and the agent needs to remember that this paddy drains faster after re-levelling, that last season AWD here cut pumping by roughly a third, and that the farmer switched from a diesel pump to an electric one in March. None of that lives in a sensor reading. It lives in past experience, and past experience needs a place to be stored and retrieved by meaning, not by primary key.
That’s what a vector database is for. And in production, I use DashVector (Alibaba Cloud’s managed vector search service) as the semantic index behind the agent’s episodic memory.
This is the post I wish I had when I started wiring it up: what a vector DB actually is, how DashVector works on Alibaba Cloud, and how I used it in a real agritech product (not a toy RAG demo).
What is a vector database?
A traditional database answers questions like “give me the row where id = 42” or “find all orders from last Tuesday”. Exact matches. Structured fields. SQL.
A vector database answers a different question: “what in my data is most similar to this?”
The trick is embeddings. You run text (or images, or code) through an embedding model and get back a fixed-length array of floats (a point in high-dimensional space). Semantically similar content ends up close together. “Farmer uses diesel pump” and “Farmer switched to electric pump” land near each other. “Paddy 3 water level dropped 4 cm overnight” lands somewhere else entirely.
A vector database stores those points and lets you query by nearest neighbours: given a new embedding, return the top-K closest documents. That’s the core of semantic search, RAG, and agent memory.
| Traditional DB | Vector DB | |
|---|---|---|
| Query style | Exact / range on fields | Similarity on embeddings |
| Best for | Transactions, profiles, sessions | Semantic recall, RAG, recommendations |
| Index | B-tree, hash | HNSW, IVF, etc. |
| Example | WHERE farmer_id = 'abc' |
“memories like this situation” |
You almost never use a vector DB alone. In every production system I’ve built (including NaLog Agent) the pattern is vector index + structured store. The vector DB holds embeddings and lightweight metadata for fast similarity search. A regular database (PostgreSQL with pgvector, or in my case Tablestore) holds the full records: timestamps, TTL, reinforcement counts, supersession links. Query the vector DB for candidates, hydrate from the structured store, rank and filter in application code.
I’ve been building RAG systems for over two years and working with vector search for more than a decade. My Vector RAG vs agentic search post covers when each retrieval strategy makes sense. DashVector is the vector side of that equation on Alibaba Cloud.
What is DashVector?
DashVector is Alibaba Cloud’s fully managed vector database. No clusters to tune, no index parameters to babysit at 3 AM. You create a collection (think: a table with a vector column), upsert documents, and query by cosine similarity over HTTP.
It sits naturally in the Alibaba Cloud AI stack:
| Service | Role in a memory/RAG pipeline |
|---|---|
| Model Studio (DashScope) | Embedding models (text-embedding-v3) and rerankers (qwen3-rerank) |
| DashVector | Vector index — top-K semantic recall |
| Tablestore | Structured storage — full memory records, TTL, profiles |
| Function Compute | Serverless app that orchestrates the pipeline |
I already use Model Studio for Qwen in Cursor and private Qwen APIs on FC. DashVector completes the picture: embeddings computed on DashScope, vectors stored on DashVector, metadata on Tablestore, all reachable from a Function Compute instance in the same region (or cross-region over HTTPS when you have to).
Why DashVector and not pgvector / Chroma / FAISS?
Honest answer: managed beats self-hosted for a serverless agent deployed on Function Compute. I don’t want a persistent vector process sitting next to my ZIP-based FC function. DashVector gives me an HTTP API, scales independently, and I pay for what I use. For local dev and tests, NaLog Agent ships a local vector driver (in-process cosine similarity over a JSON file) that mirrors the same interface — swap VECTOR_DRIVER=local to VECTOR_DRIVER=dashvector and nothing else changes.
Getting started with DashVector
1. Create a cluster
Head to the DashVector console. Create a cluster, pick a region, and note the endpoint and API key. The endpoint is cluster-specific (something like vrs-cn-xxxx.dashvector.cn-hangzhou.aliyuncs.com).
Region gotcha I hit immediately: DashVector is not available in every Alibaba Cloud region. My NaLog Agent backend runs in Thailand (ap-southeast-7, Bangkok) because that’s where the farmers are, but DashVector doesn’t exist there yet. I created the cluster in Singapore (ap-southeast-1) and reach it cross-region over HTTPS from Function Compute. Adds a few milliseconds per recall query; totally acceptable for an agent that already spends ~30 seconds on a deep reasoning turn.
2. Create a collection
A collection defines the vector dimension and distance metric. In NaLog Agent I use text-embedding-v3 at 1024 dimensions with cosine distance:
curl -X POST "https://<your-endpoint>/v1/collections" \
-H "dashvector-auth-token: <your-api-key>" \
-H "Content-Type: application/json" \
-d '{
"name": "nalog_memory",
"dimension": 1024,
"metric": "cosine",
"fields_schema": {
"farmerId": "STRING",
"paddyId": "STRING",
"memoryId": "STRING"
}
}'
The fields_schema is important. DashVector lets you attach scalar fields to each vector and filter queries on them — essential when you have thousands of memories across hundreds of farmers and only want results for one farmer at a time.
In the agent, provisioning is a single script:
VECTOR_DRIVER=dashvector npm run provision
Which calls provisionCollection() and creates the collection if it doesn’t exist.
3. Upsert documents
Each episodic memory becomes one DashVector document: the memory ID as the document ID, the embedding vector, and filter fields:
await vector.upsert(memory.memoryId, embeddingVector, {
farmerId: 'farmer_abc',
paddyId: 'paddy_3',
memoryId: memory.memoryId,
});
The embedding comes from DashScope’s text-embedding-v3 model via the OpenAI-compatible endpoint I already use everywhere else:
const res = await client.embeddings.create({
model: 'text-embedding-v3',
input: `${type}: ${text}`,
dimensions: 1024,
});
const vector = res.data[0].embedding;
Lock your embedding model. The production collection was built in text-embedding-v3’s vector space. Changing models means re-embedding every stored memory. I learned this lesson years ago with other vector stacks; DashVector doesn’t magically fix it.
4. Query by similarity
When a farmer asks a question, the agent embeds the query and asks DashVector for the top-K nearest memories, filtered to that farmer:
const hits = await vector.query(queryEmbedding, {
topK: 12,
filter: { farmerId: 'farmer_abc' },
});
DashVector’s filter syntax is SQL-like: farmerId = 'farmer_abc' AND paddyId = 'paddy_3'. Simple, effective.
The HTTP API returns hits with a score field. Here’s the gotcha that bit me during development:
DashVector returns cosine distance (lower = more similar). My local dev driver returns cosine similarity (higher = more similar).
If you rank by raw score across environments, your recall pipeline behaves differently in dev vs production. The fix: rank by position, not raw score. Take the top-K from DashVector in order, hydrate from Tablestore, then apply your own scoring blend. Don’t compare absolute numbers across drivers.
5. Delete stale vectors
When Tablestore TTL physically deletes an old memory row, the vector can become an orphan. NaLog Agent does lazy cleanup: after a query, if a vector ID has no matching Tablestore row, delete it from DashVector. Keeps the index honest without a nightly batch job.
How NaLog Agent uses DashVector
The memory system has three tiers. DashVector powers the semantic leg of episodic recall:
| Tier | Store | What it holds |
|---|---|---|
| Profile (sticky) | Tablestore | Language, irrigation style, durable preferences |
| Episodic (decaying) | Tablestore + DashVector | Dated field experience; TTL ~400 days |
| Semantic recall | DashVector | Embedding index for “what past situations resemble this one?” |
The recall pipeline
When the agent needs context, MemoryManager runs a vector-first pipeline:
- Embed the query via DashScope
text-embedding-v3. - Query DashVector for top-K candidates, filtered by
farmerId(and optionallypaddyId). - Hydrate from Tablestore with a single
BatchGetRow— point lookups, O(topK) regardless of total history size. - Rerank with DashScope’s
qwen3-rerankcross-encoder (re-orders candidates against the actual query text). - Blend four signals into a final score:
score = 0.50 × semantic_rank (reranked position)
+ 0.10 × keyword_overlap (BM25-inspired term match)
+ 0.25 × recency (half-life ≈ 120 days)
+ 0.15 × reinforcement (boosts memories reused in past turns)
- Return the top 5, summarise into a compact block, inject into the Qwen context window.
DashVector does step 2. It’s the fast filter that narrows thousands of memories down to a dozen candidates worth hydrating. Without it, you’d scan every memory for every query — fine with 50 memories, catastrophic with 50,000.
Why not just DashVector?
Because similarity ≠ truth. “Farmer uses diesel pump” and “Farmer switched to electric pump” are embedding neighbours. A pure vector search returns both and the LLM flips a coin. That’s the append-only memory problem I benchmarked: Mem0-style retrieval gets 92.9% recall but leaks 10 stale facts into the top-5.
DashVector gets you the candidates. The application layer does the hard work:
- Recency decay pushes last season’s trigger level down the ranking.
- Keyword overlap catches exact matches embeddings miss (a specific paddy name, a pump model).
- Reinforcement boosts memories the agent has actually reused.
- LLM-adjudicated supersession marks contradictions as
supersededByand excludes them from recall (but keeps them in storage for audit).
Production serves zero stale facts across 46 memories and 14 labeled queries. DashVector is necessary but not sufficient — you still need forgetting logic on top.
Swappable drivers
The vector layer is behind an interface with two implementations:
| Driver | When | How |
|---|---|---|
local |
Dev, tests, offline benchmark | In-process cosine over a JSON file |
dashvector |
Production on Alibaba Cloud | HTTP API to managed cluster |
Same upsert, query, delete methods. Same recall pipeline. I run 125 deterministic tests against the local driver and deploy with VECTOR_DRIVER=dashvector on Function Compute. The benchmark (npm run bench) reproduces the memory ablation study entirely offline.
What I learned
Vector DBs are indexes, not databases. DashVector stores vectors and a few filter fields. The real memory — text, timestamps, TTL, supersession links, reinforcement counts — lives in Tablestore. Treat DashVector like you’d treat an Elasticsearch index: fast lookup, not source of truth.
Filter fields are not optional at scale. Without farmerId filtering, a query returns the most similar memories across all farmers. Fine for a demo. Wrong for a multi-tenant agent where memory is scoped per user.
Distance vs similarity will trip you up. If you build a local dev driver (and you should), make sure your ranking logic doesn’t depend on the sign or scale of the score. Position-based ranking saved me.
Region availability matters. Plan your stack around where each service actually exists. Cross-region HTTPS between FC and DashVector works; just account for the latency and don’t assume co-location.
Embeddings are a contract. Pick a model, dimension, and metric at provisioning time and stick with them. Re-embedding everything is a migration, not a config change.
Forgetting is the product feature. A vector DB that only grows is a vector DB that eventually serves stale, contradictory noise. TTL on the structured store, recency decay in ranking, and explicit supersession for contradictions — DashVector handles the “find similar” part; you handle the “forget what’s no longer true” part.
Try it yourself
The full implementation is open source: github.com/khawtech/nalog-agent. Docker quick-start with demo data and only a DashScope API key. Flip VECTOR_DRIVER=dashvector and STORAGE_DRIVER=alibaba, run npm run provision, and you have the same production stack I described here.
If you want the bigger picture — the Qwen models, the ReAct loop, the human-in-the-loop pump approval, the benchmark numbers — read the full NaLog Agent write-up. This post is the DashVector chapter.
And if you’re new to Alibaba Cloud’s AI stack, start with testing Model Studio and running Qwen in Cursor. DashVector is the piece that turns those embeddings into something an agent can remember.