Skip to content

Commit c9d73bb

Browse files
authored
fix(knowledge): walk a large bounded set on the row before ranking it exactly (#8106)
* fix(knowledge): walk a large bounded set on the row before ranking it exactly A member reading most of a large source, with that source selected as a filter, enumerated a bounded permitted set of tens of thousands of documents and then ranked every chunk of it exactly on both legs: the vector leg read every chunk's projected vector, and the keyword leg materialized every chunk of the set before it matched the term. Cold, each leg outran its budget and the search returned nothing. A bounded set past a size limit is now walked on the row first, where the plan's source and ACL decide readability and the walk stops at its tuple cap, and ranked exactly only when the walk cannot fill its pool, so recall is never below the exact ranking's. The keyword leg treats the same set as a narrow on-row reader: Tin windows where Tin serves, otherwise the GIN shape whose cost follows the term's matches. Sets under the limit keep their exact paths. * fix(knowledge): refill a large bounded set's pool with the exact ranking once hydration runs it short The walk decides readability on the projection row, which is broader than the document predicate hydration applies, so a pool the walk filled can still run short of readable rows. The refill for a large bounded set is now the exact ranking, complete over the set, placed behind the rows already read so the pages keep their offsets. * fix(knowledge): rank a large bounded set's refill past the rows already read The refill's exact ranking excludes the chunks the pool already holds inside the statement, so every refill is a full window of fresh rows rather than a window thinned by the rows the walk found first. * fix(knowledge): hand a large bounded set's exhausted Tin windows to the GIN ranking A narrow reader's page is left short by design once the widest window cannot fill it; a large bounded set's read was exhaustive before, so its widest window that still falls short now hands the page to the GIN ranking, which covers every match. * fix(knowledge): hand only a large bounded set's first page to the GIN ranking Tin and GIN order candidates differently, so an offset advanced through one ranking cannot resume the other. A large bounded set's first page that Tin's widest window cannot fill goes to GIN; a later page stays with Tin and is left short as a narrow reader's is.
1 parent 98db661 commit c9d73bb

2 files changed

Lines changed: 268 additions & 54 deletions

File tree

‎apps/sim/lib/knowledge/search/queries.test.ts‎

Lines changed: 149 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -41,6 +41,7 @@ import {
4141
handleTagAndVectorSearch,
4242
handleTagOnlySearch,
4343
handleVectorOnlySearch,
44+
PERMITTED_EXACT_DOCUMENT_LIMIT,
4445
type PermittedDocuments,
4546
resolvePermittedDocuments,
4647
resolveReach,
@@ -1794,6 +1795,86 @@ describe('permitted-document planner', () => {
17941795
expect(ginStatements()).toHaveLength(0)
17951796
})
17961797

1798+
describe('a bounded set past the exact-ranking size', () => {
1799+
const large = Array.from({ length: PERMITTED_EXACT_DOCUMENT_LIMIT }, (_, index) => ({
1800+
id: `doc-${index}`,
1801+
connectorId: 'src-a',
1802+
}))
1803+
const accessPlan = {
1804+
connectors: { workspace: [], admin: ['src-a'], members: [], liveProofRequired: [] },
1805+
observers: { confirmed: [], observed: [] },
1806+
memberSources: [],
1807+
connectorTypes: new Map(),
1808+
uploads: true,
1809+
}
1810+
1811+
it('ranks with Tin as a narrow reader, decided on the row', async () => {
1812+
tinPages = [{ ranked: 1500, candidates: [hit('a', 'src-a')] }]
1813+
queueTableRows(schemaMock.embedding, [{ ...hit('a', 'src-a'), content: 'release notes' }])
1814+
const results = await keyword({
1815+
permitted: { kind: 'bounded', documents: large },
1816+
accessPlan,
1817+
})
1818+
expect(results.map((row) => row.id)).toEqual(['a'])
1819+
expect(mockResolveTinKeywordQuery).toHaveBeenCalledTimes(1)
1820+
expect(tinStatements()).toHaveLength(1)
1821+
expect(JSON.stringify(tinStatements()[0])).toContain('2000')
1822+
expect(JSON.stringify(tinStatements()[0])).not.toContain('doc-4999')
1823+
expect(ginStatements()).toHaveLength(0)
1824+
})
1825+
1826+
it('falls back to a GIN ranking that reads what the term matches, not every chunk of the set', async () => {
1827+
mockResolveTinKeywordQuery.mockResolvedValue(null)
1828+
queueTableRows(schemaMock.embedding, [{ ...hit('a', 'src-a'), content: 'release notes' }])
1829+
await keyword({ permitted: { kind: 'bounded', documents: large }, accessPlan })
1830+
expect(tinStatements()).toHaveLength(0)
1831+
expect(ginStatements()).toHaveLength(1)
1832+
expect(JSON.stringify(ginStatements()[0])).not.toContain('doc-4999')
1833+
})
1834+
1835+
it('hands the page to the GIN ranking when the widest window cannot fill it', async () => {
1836+
tinPages = [
1837+
{ ranked: 2000, candidates: [] },
1838+
{ ranked: 20_000, candidates: [] },
1839+
]
1840+
queueTableRows(schemaMock.embedding, [{ ...hit('a', 'src-a'), content: 'release notes' }])
1841+
await keyword({ permitted: { kind: 'bounded', documents: large }, accessPlan })
1842+
expect(tinStatements()).toHaveLength(2)
1843+
/** Every match is covered again, by the ranking whose cost follows the term, not the set. */
1844+
expect(ginStatements()).toHaveLength(1)
1845+
expect(JSON.stringify(ginStatements()[0])).not.toContain('doc-4999')
1846+
})
1847+
1848+
it('leaves a later page short rather than resuming a different ranking at its offset', async () => {
1849+
/** The first page fills from Tin; hydration keeps half, so a second page is asked for. */
1850+
const first = Array.from({ length: 40 }, (_, index) => hit(`t-${index}`, 'src-a'))
1851+
tinPages = [
1852+
{ ranked: 2000, candidates: first },
1853+
{ ranked: 2000, candidates: [] },
1854+
{ ranked: 20_000, candidates: [] },
1855+
]
1856+
queueTableRows(
1857+
schemaMock.embedding,
1858+
first.slice(0, 20).map((row) => ({ ...row, content: 'release notes' }))
1859+
)
1860+
const results = await keyword({
1861+
topK: 40,
1862+
permitted: { kind: 'bounded', documents: large },
1863+
accessPlan,
1864+
})
1865+
expect(results).toHaveLength(20)
1866+
expect(tinStatements()).toHaveLength(3)
1867+
expect(ginStatements()).toHaveLength(0)
1868+
})
1869+
1870+
it('keeps the bounded read for a set under the size', async () => {
1871+
mockResolveTinKeywordQuery.mockResolvedValue(null)
1872+
await keyword({ permitted: { kind: 'bounded', documents: large.slice(0, -1) }, accessPlan })
1873+
expect(mockResolveTinKeywordQuery).not.toHaveBeenCalled()
1874+
expect(JSON.stringify(ginStatements()[0])).toContain('doc-4998')
1875+
})
1876+
})
1877+
17971878
it('leaves the page to the GIN ranking when the widest window cannot fill it', async () => {
17981879
tinPages = [
17991880
{ ranked: 2000, candidates: [] },
@@ -2400,6 +2481,74 @@ describe('filters on a resolved scope', () => {
24002481
expect(JSON.stringify(walks[0])).toContain('release')
24012482
})
24022483

2484+
describe('a bounded set past the exact-ranking size', () => {
2485+
const large = Array.from({ length: PERMITTED_EXACT_DOCUMENT_LIMIT }, (_, index) => ({
2486+
id: `doc-${index}`,
2487+
connectorId: 'src-a',
2488+
}))
2489+
const walked = Array.from({ length: 200 }, (_, index) => hit(`w-${index}`, 'src-a'))
2490+
const search = (documents: typeof large) =>
2491+
handleVectorOnlySearch({
2492+
...params,
2493+
permitted: { kind: 'bounded', documents },
2494+
accessPlan: plan(),
2495+
})
2496+
beforeEach(() => {
2497+
const execute = dbChainMockFns.execute.getMockImplementation()!
2498+
dbChainMockFns.execute.mockImplementation(async (query) => {
2499+
/** The projection is filled, so a walk decides readability on the row. */
2500+
if (render(query).sql.includes('AS unfilled')) return [{ unfilled: false }]
2501+
return execute(query)
2502+
})
2503+
})
2504+
2505+
it('walks the graph on the row instead of ranking every chunk of the set', async () => {
2506+
traversedRows = walked
2507+
queueTableRows(schemaMock.embedding, [walked[0]])
2508+
expect((await search(large)).map((row) => row.id)).toEqual(['w-0'])
2509+
const walks = statements().filter((query) => isWalk(query.sql))
2510+
expect(walks).toHaveLength(1)
2511+
expect(statements().filter((query) => isExactRanking(query.sql))).toHaveLength(0)
2512+
/** Readability rides on the row through the plan; the set's identifiers never cross the wire. */
2513+
expect(JSON.stringify(walks[0])).not.toContain('doc-4999')
2514+
expect(JSON.stringify(walks[0])).toContain('src-a')
2515+
})
2516+
2517+
it('ranks the set exactly when the walk cannot fill its pool', async () => {
2518+
traversedRows = []
2519+
exactRows = [hit('a', 'src-a')]
2520+
queueTableRows(schemaMock.embedding, [hit('a', 'src-a')])
2521+
expect((await search(large)).map((row) => row.id)).toEqual(['a'])
2522+
expect(statements().filter((query) => isWalk(query.sql))).toHaveLength(1)
2523+
const exact = statements().filter((query) => isExactRanking(query.sql))
2524+
expect(exact).toHaveLength(1)
2525+
expect(JSON.stringify(exact[0])).toContain('doc-4999')
2526+
})
2527+
2528+
it('refills a pool the walk filled but hydration could not with the exact ranking', async () => {
2529+
traversedRows = walked
2530+
exactRows = [hit('a', 'src-a')]
2531+
/** None of the walked rows survives the document predicate; the set's own ranking then does. */
2532+
for (let page = 0; page < 10; page++) queueTableRows(schemaMock.embedding, [])
2533+
queueTableRows(schemaMock.embedding, [hit('a', 'src-a')])
2534+
expect((await search(large)).map((row) => row.id)).toEqual(['a'])
2535+
expect(statements().filter((query) => isWalk(query.sql))).toHaveLength(1)
2536+
const exact = statements().filter((query) => isExactRanking(query.sql))
2537+
expect(exact).toHaveLength(1)
2538+
/** The refill ranks past the rows already read, so nothing already rejected is read twice. */
2539+
expect(JSON.stringify(exact[0])).toContain('doc-4999')
2540+
expect(JSON.stringify(exact[0])).toContain('w-199')
2541+
})
2542+
2543+
it('ranks a set under the size exactly, without a walk', async () => {
2544+
exactRows = [hit('a', 'src-a')]
2545+
queueTableRows(schemaMock.embedding, [hit('a', 'src-a')])
2546+
expect((await search(large.slice(0, -1))).map((row) => row.id)).toEqual(['a'])
2547+
expect(statements().filter((query) => isWalk(query.sql))).toHaveLength(0)
2548+
expect(statements().filter((query) => isExactRanking(query.sql))).toHaveLength(1)
2549+
})
2550+
})
2551+
24032552
it('ranks a date-bounded set exactly even when a member source has its own index', async () => {
24042553
indexedSourceRows = [{ name: 'idx', connectorId: 'member-src' }]
24052554
exactRows = [{ id: 'a' }]

0 commit comments

Comments
 (0)