Three weeks ago you saved an article about hydrostatic pressure in parking garages. Today you need it, and all you can recall is "the water pressure thing". What do you type? Keyword search fails here because the query and the document share no words, a failure mode retrieval people call vocabulary mismatch. Stashio ships semantic search for exactly this case. The product story lives in the second-brain post: my pre-Stashio Apple Notes held roughly 2,400 unsorted links, and after three months of daily dogfooding about 94% of my retrieval attempts succeed in under 10 s. This post is the machinery behind the semantic half of that number.
The constraint, as always at this studio, is that the machinery runs on the phone. No embedding API, no vector database in someone's data center, no query leaving the device. That rules out most of the standard retrieval stack, and, it turns out, none of the useful parts.
From sentences to vectors
A sentence embedding model reads a piece of text and outputs a fixed-length vector. Ours outputs 384 numbers. Training pushes texts with similar meaning toward nearby points in that 384-dimensional space, so "hydrostatic pressure in a basement garage" and "water pressure underground parking" land close together even though they share one word. The geometry does the matching that the vocabulary cannot.
Stashio uses a distilled MiniLM-class sentence encoder with 384-dimensional output, weights quantized to int8, about 25 MB inside the app bundle. It is a multilingual model because our users mix Vietnamese and English in the same library, sometimes in the same query. A multilingual embedding space puts "áp suất thủy tĩnh" near "hydrostatic pressure", which means a Vietnamese query finds an English article without any translation step. On iOS the encoder runs through Core ML, on Android through LiteRT. Studio bench note: median encode time for a title-plus-excerpt input is 31 ms on an iPhone 15 with the Neural Engine and 54 ms on a Pixel 8 on CPU, 200 runs per device, release builds, batteries above 50%.
Cosine similarity, normalized early
Two vectors are compared by the angle between them: cos θ = a·b / (|a||b|). Identical direction scores 1, unrelated directions score near 0. Raw vector magnitude mostly encodes text length and token frequency artifacts, so we throw it away by L2-normalizing every vector once, at save time. After normalization |a| = |b| = 1 and cosine collapses to a plain dot product, 384 multiplies and adds. That is the whole trick. For unit vectors Euclidean distance ranks identically anyway, since |a − b|² = 2 − 2·(a·b), so nothing is lost by picking the cheaper formula.
Brute force is fine below 10k items
Search is a loop. Embed the query, dot it against every stored vector, keep the top 20 in a small heap. Over 10,000 items that is 3.84 million multiply-adds per query, which sounds expensive until you measure it. Studio bench note: 10,000 synthetic 384-dim vectors, 200 queries, medians. The float32 scan takes 1.8 ms on an iPhone 15 with vDSP and 3.4 ms on a Pixel 8 with NEON intrinsics. The int8 scan takes 0.9 ms and 1.6 ms on the same two devices. The encoder needs 31 ms just to embed the query. Next to the encoder, the scan barely registers.
Approximate indexes like HNSW and IVF exist for a real reason, and that reason is a hundred million vectors in a server rack. On a phone they charge rent: roughly 1.5× memory for graph links, insert work on every share-sheet save, tombstone bookkeeping when you delete, recall that drops below 1.0, and three tuning parameters per index. The exact scan has recall 1.0 by construction, zero tuning, and delete is a row delete. On our own numbers the crossover where an approximate index starts paying for itself sits somewhere above 50,000 vectors. My whole library after four years of hoarding is 2,400 items, about 7,900 vectors once the PDFs are chunked. There is an HNSW branch in the repo. It has never been merged.
Int8 and the price of a byte
A 384-dim float32 vector costs 1,536 bytes. At 10,000 items the index weighs 15.4 MB, which a modern phone shrugs at in storage but feels in memory bandwidth during a scan. Quantizing each vector to int8 cuts it to 384 bytes plus one 4-byte scale factor, 3.9 MB for the same 10,000 items, and the narrower reads are the main reason the int8 scan above runs twice as fast.
The scheme is symmetric per-vector quantization. Take the normalized vector, compute s = max|x_i| / 127, store q_i = round(x_i / s) together with s itself. A dot product then runs in int32 accumulators and rescales once at the end: a·b ≈ s_a·s_b·Σ q_a[i]·q_b[i]. ARM's sdot instruction eats four of those multiply-adds per lane per cycle.
// per query: embed, normalize, quantize, then scan
fun topK(q: QVec, items: List<QVec>, k: Int): List<Hit> {
val heap = BoundedMinHeap<Hit>(k)
for (item in items) {
var acc = 0 // int32
for (i in 0 until 384) { // NEON sdot in practice
acc += q.q[i] * item.q[i]
}
val score = acc * q.scale * item.scale
heap.offer(Hit(item.id, score))
}
return heap.sortedByDescending { it.score }
}
Quantization costs accuracy, so we measured the cost. Against exact float32 top-10 lists on my dogfood library, replaying 500 real queries from my own local search history, int8 reproduced 98.4% of the results. Every disagreement sat in ranks 8 through 10, where neighboring scores differ by less than 0.004 and the ordering is honestly arbitrary. For a bookmark app that trade is free money. One detail matters: quantize the normalized vector. Quantizing first and normalizing after wastes int8 range on magnitude you were about to divide out.
Chunking long PDFs
One vector per item works for links and short notes. It fails for a 68-page motor datasheet, because the encoder reads about 256 tokens and a single vector for 68 pages averages everything into mud. Stashio splits extracted PDF text into windows of roughly 180 words with a one-sentence overlap, so no idea gets cut mid-thought. Each window becomes its own vector, tagged with the item id and the page number. That datasheet becomes 214 chunks, about 83 KB of int8 index; opening the result deep-links to the matching page.
An item's score is the maximum over its chunk scores. We tried mean pooling first and discarded it, because averaging punishes long documents: one perfect page drowns among 213 mediocre ones. Max pooling scores a document by its best page, which matches how people remember PDFs anyway, by the one diagram they need.
Hybrid ranking, because keywords still win exact matches
Embeddings are terrible at identifiers. The query "E9" should return the screenshot of pillar E9 on floor B2 from my garage logs, and no 384-dim vector reliably separates E9 from E7. Part numbers, error codes, ISO standard names, all the same story. So keyword search stays: SQLite FTS5 with BM25 scoring over titles, tags and extracted text. Each engine covers the other's blind spot, paraphrase on one side and exact tokens on the other.
Merging two ranked lists is its own small science. Our first attempt was a weighted sum of normalized scores, discarded after a week, because BM25 scores and cosine scores live on unrelated scales and every normalization we tried was fragile against outlier queries. What shipped is reciprocal rank fusion: score(d) = Σ 1/(60 + rank_i(d)), summed over both lists. Only ranks matter, the scales cancel, and the constant 60 keeps a single first-place vote from steamrolling an item both engines ranked fifth. It is one line of code. It has survived every query type we have thrown at it.
Syncing an index the server cannot read
Published inversion results keep showing that sentence embeddings can be decoded back into much of their source text. A vector is content, not metadata, and Stashio treats it exactly like the bookmarks themselves. When you enable the optional account, the index syncs through atuan's backend as encrypted blobs: shards of a few hundred vectors each, encrypted on the device before upload, with keys that never leave your hardware. The server stores ciphertext and a version counter per shard. It can tell how big your library is and when it changed. It cannot search it, and neither can we.
Merging across devices happens locally after download, per item id, newest write wins. Each vector also carries the encoder model version, so when we ship a better model the app re-embeds stale items lazily in the background and a mixed-version library stays searchable the whole time. Because the index is built on the phone before any of this, search works in airplane mode, in a basement garage, on an install with no account at all. The reasoning behind that default is the studio's on-device-first post.
The latency budget, honestly
Add it up: 31 ms to embed the query, about 1 ms to scan, under 1 ms for FTS5 plus fusion over a few thousand rows. Call it 40 ms from keystroke to ranked results, or 70 ms on the slower bench phone. The 94%-of-retrievals-under-10 s figure from the top of this post was never limited by compute; the other 9.9 s is a human remembering that the thing they want had something to do with water pressure. The phone's job is to be ready the moment the words arrive, with no server in the loop and nothing to be down. The short pitch, screenshots included, is on the Stashio app page.
Ba tuần trước bạn lưu một bài về áp suất thủy tĩnh trong hầm đỗ xe. Hôm nay cần đến, trong đầu chỉ còn "cái vụ áp lực nước". Gõ gì bây giờ? Tìm theo keyword thua ngay từ vòng gửi xe vì query và tài liệu không chung từ nào. Dân retrieval gọi đây là vocabulary mismatch. Stashio làm semantic search chính vì tình huống này. Chuyện sản phẩm đã kể trong bài second brain: Apple Notes cũ của mình chứa khoảng 2.400 link không sort; sau ba tháng dogfood hằng ngày, khoảng 94% lần tìm của mình thành công dưới 10 giây. Bài này là phần máy móc đứng sau nửa semantic của con số đó.
Ràng buộc quen thuộc của studio: máy móc chạy trên điện thoại. Không API embedding, không vector database trong data center của ai đó, không query nào rời máy. Nghe như mất gần hết stack retrieval tiêu chuẩn. Thực tế chẳng mất phần nào hữu ích.
Từ câu chữ thành vector
Model sentence embedding đọc một đoạn text rồi trả về vector độ dài cố định. Bản của Stashio trả 384 số. Huấn luyện đẩy các đoạn text gần nghĩa về các điểm gần nhau trong không gian 384 chiều, nên "áp suất thủy tĩnh hầm xe" và "áp lực nước bãi đỗ ngầm" nằm sát nhau dù chỉ chung một từ. Hình học làm nốt phần việc mà từ vựng bó tay.
Stashio dùng encoder loại MiniLM distilled, output 384 chiều, trọng số quantize về int8, chiếm khoảng 25 MB trong bundle app. Model đa ngôn ngữ vì user của tụi mình trộn tiếng Việt lẫn tiếng Anh trong cùng thư viện, có khi trong cùng một query. Không gian embedding đa ngôn ngữ đặt "áp suất thủy tĩnh" cạnh "hydrostatic pressure", nghĩa là query tiếng Việt tìm ra bài tiếng Anh mà không cần bước dịch nào. Trên iOS encoder chạy qua Core ML, trên Android qua LiteRT. Ghi chú bench của studio: thời gian encode trung vị cho input gồm title cộng excerpt là 31 ms trên iPhone 15 dùng Neural Engine và 54 ms trên Pixel 8 chạy CPU, mỗi máy 200 lần, bản release, pin trên 50%.
Cosine và chuyện chuẩn hóa sớm
So hai vector bằng góc giữa chúng: cos θ = a·b / (|a||b|). Cùng hướng thì điểm 1. Không liên quan thì gần 0. Độ lớn thô của vector chủ yếu mang artifact về độ dài text và tần suất token, nên bỏ luôn: chuẩn hóa L2 mỗi vector đúng một lần, ngay lúc lưu. Sau đó |a| = |b| = 1 và cosine rút gọn thành phép dot product trần, 384 phép nhân cộng. Toàn bộ mẹo nằm ở đó. Với vector đơn vị, khoảng cách Euclid xếp hạng y hệt vì |a − b|² = 2 − 2·(a·b), nên chọn công thức rẻ hơn không mất gì.
Dưới 10 nghìn item cứ quét thẳng
Search là một vòng lặp. Embed query, dot với từng vector đã lưu, giữ top 20 trong heap nhỏ. Với 10.000 item là 3,84 triệu phép nhân cộng mỗi query. Nghe nặng cho tới khi đo. Ghi chú bench: 10.000 vector 384 chiều sinh ngẫu nhiên, 200 query, lấy trung vị. Quét float32 hết 1,8 ms trên iPhone 15 với vDSP, 3,4 ms trên Pixel 8 với NEON. Quét int8 hết 0,9 ms và 1,6 ms trên cùng hai máy. Riêng encoder đã cần 31 ms để embed query. Đứng cạnh encoder, vòng quét gần như không đáng kể.
Index xấp xỉ như HNSW hay IVF sinh ra có lý do. Lý do đó là trăm triệu vector trong rack server. Trên điện thoại chúng thu tiền thuê: thêm chừng 1,5 lần bộ nhớ cho link đồ thị, tốn công insert mỗi lần lưu từ share sheet, dọn tombstone khi xóa, recall tụt dưới 1,0, thêm ba tham số phải tune cho mỗi index. Quét chính xác có recall 1,0 theo định nghĩa, khỏi tune, xóa là xóa một row. Theo số đo của tụi mình, điểm hòa vốn của index xấp xỉ nằm đâu đó trên 50.000 vector. Cả thư viện của mình sau bốn năm tích trữ mới 2.400 item, cỡ 7.900 vector sau khi chặt PDF thành chunk. Trong repo có một branch HNSW. Chưa bao giờ được merge.
Int8 và cái giá của từng byte
Một vector float32 384 chiều tốn 1.536 byte. Với 10.000 item, index nặng 15,4 MB. Lưu trữ thì điện thoại chấp, nhưng băng thông bộ nhớ lúc quét thì cảm nhận rõ. Quantize mỗi vector về int8 còn 384 byte cộng một hệ số scale 4 byte, tức 3,9 MB cho cùng 10.000 item. Đọc ít byte hơn là lý do chính vòng quét int8 ở trên nhanh gấp đôi.
Cách làm là quantize đối xứng theo từng vector. Lấy vector đã chuẩn hóa, tính s = max|x_i| / 127, lưu q_i = round(x_i / s) kèm chính s. Dot product chạy trong bộ tích lũy int32 rồi nhân lại một lần cuối: a·b ≈ s_a·s_b·Σ q_a[i]·q_b[i]. Lệnh sdot của ARM nuốt bốn phép nhân cộng mỗi lane mỗi chu kỳ.
// mỗi query: embed, chuẩn hóa, quantize, rồi quét
fun topK(q: QVec, items: List<QVec>, k: Int): List<Hit> {
val heap = BoundedMinHeap<Hit>(k)
for (item in items) {
var acc = 0 // int32
for (i in 0 until 384) { // thực tế là NEON sdot
acc += q.q[i] * item.q[i]
}
val score = acc * q.scale * item.scale
heap.offer(Hit(item.id, score))
}
return heap.sortedByDescending { it.score }
}
Quantize thì mất độ chính xác, nên phải đo. So với top-10 float32 chính xác trên thư viện dogfood của mình, replay 500 query thật lấy từ lịch sử tìm kiếm local, int8 tái hiện 98,4% kết quả. Mọi chỗ lệch đều nằm ở hạng 8 tới 10, nơi điểm số cạnh nhau chênh chưa tới 0,004 và thứ tự thật ra ngẫu nhiên. Với app bookmark, đổi vậy là lời to. Một chi tiết quan trọng: quantize vector đã chuẩn hóa. Quantize trước rồi mới chuẩn hóa là phí dải int8 cho phần độ lớn sắp bị chia bỏ.
Chặt PDF dài thành chunk
Mỗi item một vector là đủ cho link và note ngắn. Với datasheet motor 68 trang thì hỏng, vì encoder chỉ đọc chừng 256 token và một vector cho 68 trang là nghiền tất cả thành bùn. Stashio cắt text trích từ PDF thành các cửa sổ chừng 180 từ, chồng lấn một câu để không ý nào bị đứt giữa chừng. Mỗi cửa sổ thành một vector riêng, gắn id item cùng số trang. Cuốn datasheet kia thành 214 chunk, cỡ 83 KB index int8; mở kết quả là nhảy thẳng tới đúng trang.
Điểm của item lấy max trên các chunk của nó. Ban đầu thử mean pooling rồi bỏ: lấy trung bình là phạt tài liệu dài, vì một trang khớp hoàn hảo chìm giữa 213 trang thường thường. Max pooling chấm tài liệu bằng trang tốt nhất, đúng kiểu người ta nhớ PDF: nhớ mỗi cái sơ đồ mình cần.
Xếp hạng lai vì keyword vẫn thắng khoản khớp chính xác
Embedding rất tệ với mã định danh. Query "E9" phải trả về tấm screenshot cột E9 tầng B2 trong log hầm xe của mình, mà chẳng vector 384 chiều nào tách nổi E9 với E7 một cách đáng tin. Mã linh kiện, mã lỗi, tên tiêu chuẩn ISO, chung số phận. Nên keyword search ở lại: SQLite FTS5 chấm điểm BM25 trên title, tag cùng text trích xuất. Hai engine che điểm mù cho nhau, một bên lo diễn đạt khác lời, bên kia lo token chính xác.
Trộn hai danh sách xếp hạng là một môn khoa học nhỏ. Thử đầu tiên là cộng có trọng số các điểm đã chuẩn hóa, bỏ sau một tuần, vì điểm BM25 với điểm cosine sống trên hai thang không liên quan và mọi cách chuẩn hóa đều mong manh trước query lạ. Bản ship dùng reciprocal rank fusion: score(d) = Σ 1/(60 + rank_i(d)), cộng trên cả hai danh sách. Chỉ hạng có nghĩa, thang điểm tự triệt tiêu, còn hằng số 60 giữ cho một phiếu hạng nhất không đè bẹp item được cả hai engine xếp hạng năm. Một dòng code. Tới giờ chưa thua kiểu query nào.
Sync một index mà server không đọc được
Các kết quả inversion đã công bố cứ lặp lại một điều: từ sentence embedding có thể giải ngược ra phần lớn text gốc. Vector chính là nội dung, nên được đối xử hệt như bản thân bookmark. Khi bật tài khoản tùy chọn của Stashio, index sync qua backend của atuan dưới dạng blob mã hóa: mỗi shard vài trăm vector, mã hóa ngay trên máy trước khi upload, khóa không bao giờ rời phần cứng của bạn. Server giữ ciphertext cùng một bộ đếm version cho mỗi shard. Nó biết thư viện bạn to cỡ nào và đổi lúc nào. Nó không search được, tụi mình cũng không.
Gộp dữ liệu giữa các máy diễn ra local sau khi tải về, theo id item, bản ghi mới hơn thắng. Mỗi vector còn mang version của model encoder, nên khi studio ship model tốt hơn, app re-embed dần các item cũ ở background và thư viện lẫn hai version vẫn search bình thường suốt quá trình. Vì index dựng trên điện thoại trước khi có mọi thứ trên, search chạy ở chế độ máy bay, dưới hầm B2, trên máy chưa từng đăng nhập. Lý do đứng sau mặc định này nằm ở bài on-device-first.
Định mức độ trễ, nói thẳng
Cộng lại: 31 ms embed query, chừng 1 ms quét, dưới 1 ms cho FTS5 cộng fusion trên vài nghìn row. Tính tròn 40 ms từ lúc gõ phím tới danh sách kết quả, hoặc 70 ms trên máy bench chậm hơn. Con số 94% lần tìm dưới 10 giây ở đầu bài chưa bao giờ nghẽn ở compute; 9,9 giây còn lại là con người đang nhớ ra thứ mình cần dính dáng tới áp lực nước. Việc của điện thoại là sẵn sàng đúng khoảnh khắc chữ hiện ra trong đầu bạn, không server nào trong vòng lặp, không có gì để sập. Phần giới thiệu ngắn kèm screenshot nằm ở trang app Stashio.