Repository navigation
feat(native): keep conversation history within a token budget - #79
Conversation
ChatViewModel sent the last 6 messages verbatim. With detailed answers that reached ~1,700 tokens, read again for every question, since the history comes after the retrieved sources in the prompt. HistoryBudget ports src/routing/historyBudget.ts from rferrari#72 (brought here with its vitest tests): newest turns first within 400 tokens, questions whole, earlier answers cut to their opening sentences (80 tokens), the latest exchange always kept. Used for the answer's history and for the summary's input. Golden test against the TypeScript: scripts/export-native-golden-history.mjs (438 leadSentences cases on corpus passages and edge cases, 34 conversations at two budgets).
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. Please check the settings in the CodeRabbit UI or the ⚙️ Run configuration
You can disable this status message by setting the Use the checkbox below for a quick retry:
✨ Finishing Touches 💡 1
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
Approve. The golden file regenerates byte for byte from the TS, and HistoryBudgetGoldenTest |
|
@coderabbitai review |
|
|
@coderabbitai review |
|
|
@coderabbitai review |
|
Keep conversation history within a token budget
Builds on #76 (CI); the last commit is this PR's. Finding 5 of the review.
ChatViewModelsent the last 6 messages verbatim (VERBATIM_MESSAGE_COUNT). With detailed answers (~500 tokens each) that reached ~1,700 tokens. The history sits after the retrieved sources in the prompt (RagPure.assembleChatMessages), and the sources change every question, so all of it was read again for every question. In the Expo app the same history took a follow-up's first word from 1.6 s (new chat) to 20–48 s (Pixel 6a, 1.5B, Oct 3).HistoryBudget(engine,answer/) portssrc/routing/historyBudget.tsfrom #72:It's used in two places in
ChatViewModel: the answer's history, and the summary's input (the older turns, which could be thousands of tokens).Golden test.
src/routing/historyBudget.tsand its vitest tests come from #72 unchanged, sincefullnative-devdoesn't have them.scripts/export-native-golden-history.mjsruns the real TS and writesgolden/history-budget.json:leadSentencescases: 60 corpus passages and 13 edge cases (abbreviations, decimals, quotes, newlines, Portuguese, one huge sentence), each at caps of 0–200 tokens;budgetHistorycases: 17 conversations at the default and at a tight budget, 8 of which drop turns.HistoryBudgetGoldenTestchecks the Kotlin output matches exactly. The Kotlin usesCompress.splitSentencesandCompress.approxTokens, which port the same functions (thesplitSentencesinsrc/rag/compress.tsandsrc/routing/context.tsare character-for-character identical, asResearchContextalready relies on).Not run locally. No Android toolchain here, so CI (#76) is the first place the Kotlin compiles and the golden test runs; the TS tests pass locally (7/7). Not measured on a phone. No CodeRabbit CLI review was run before pushing (
crisn't installed here).Touches
ChatViewModel.ktat the history and summary lines only, not atsend()'s start (#78) orpreload().