# Nikita Pochaev -- Full Site Content > AI Engineer and Senior Mobile Developer specializing in Android, Kotlin Multiplatform Mobile (KMM), and ML-powered product features. Recognized for shipping production-grade mobile applications, building release automation pipelines, and integrating AI/ML into mobile products. ## About Nikita Pochaev Nikita Pochaev (@po4yka) is an AI Engineer and Senior Mobile Developer with deep expertise across the Android and iOS ecosystems. He architects scalable Kotlin Multiplatform Mobile (KMM) solutions that achieve 60%+ code sharing while preserving fully native user experiences on each platform — and brings the same engineering discipline to integrating ML-powered features into production mobile products. Nikita is recognized for his expertise in MobileOps -- building release automation pipelines, optimizing Gradle build systems, and designing CI/CD infrastructure that dramatically reduces deploy times. He has led platform teams at scale, migrating large codebases (80+ screens) from legacy XML views to Jetpack Compose, reducing CI build times from 18 to 7 minutes through caching and parallelization strategies, and mentoring developers through career advancement. ### Areas of expertise - AI/ML integration -- embedding LLM-powered flows and on-device inference into native mobile experiences, bridging product intent and model capabilities - Kotlin Multiplatform Mobile (KMM) architecture -- shared networking, persistence, and domain logic across Android and iOS - Jetpack Compose and SwiftUI -- expert-level modern declarative UI, including Compose compiler stability optimization - MobileOps and CI/CD -- Gradle build optimization, Fastlane distribution, release automation pipelines - Developer tooling -- Gradle plugins, structured logging libraries, compiler metrics dashboards - Mobile architecture -- offline-first patterns, modularization strategy, clean architecture at scale ### Notable achievements - Designed shared KMP modules for multiple production projects, achieving 60%+ code sharing - Built release automation pipelines that reduced deploy time by 70% - Migrated 80+ screens from XML to Jetpack Compose - Reduced CI build times from 18 minutes to 7 minutes - Consulted on modularization strategy for a 500k LOC Android codebase - Published open-source Android libraries with community adoption (200+ stars) - Contributed to AndroidX and community libraries --- ## Blog Posts ### The Network Isn't Broken Everywhere, Just Here: Diagnosing One Connection at a Time Published: Jun 2026 | Category: Networking | Tags: DPI, Network Diagnostics, Rust, Android Over mobile data, Telegram won't bring up a connection: the client sits on "Connecting…". Over Wi-Fi, the same client connects instantly. Between the two attempts only the network changes: the point of attachment and the carrier. That is enough for traffic to the same server to go through on one network and die on the handshake on the other. The network itself is nominally fine: DNS resolves, other sites load, a ping to 8.8.8.8 comes back clean. What drops is specific connections, and always the same ones. This is the signature of DPI: a box on the carrier side inspects the traffic and drops the connections that match a signature. The rest goes through untouched. On carrier networks this asymmetry has long been the norm. A mobile carrier doesn't stop at routing: it fingerprints TLS and QUIC handshakes, throttles per connection, clamps the MTU, and strips ECN; an in-path box kills a connection that a home router would pass without a second look. So one destination is dead, the next one is fine, and any global "turn it on everywhere" switch is wrong for at least one of them. Most tools in this space start from a guess. They run one packet-level trick against a list of hosts and see whether it lands. Route-everything tunnels go to the opposite extreme: they funnel all traffic through a remote node and pay for it in latency even where nothing was broken. Both prescribe the treatment without making a diagnosis. RIPDPI runs the diagnosis first. ## Diagnose before fixing Diagnosis runs as a chain of four Rust crates: `ripdpi-diagnostics-candidates` builds the probe inputs, `ripdpi-diagnostics-probes` defines the `Probe` trait every check implements, `ripdpi-diagnostics-classification` turns raw observations into a verdict, and `ripdpi-diagnostics-runner` drives the whole battery. More than a dozen probe stages: DNS integrity and tampering, domain and QUIC reachability, an ECH handshake check, MTProto reachability for Telegram, throughput, a DoH-JSON resolver survey. Not one talks to a hard-coded server: the target arrives at runtime through a `ProbeContext`, and the probe hits exactly the address being opened. `TcpRunner` opens one TLS session and sends up to 16 padded HTTP HEAD requests, each larger than the last, tracking the cumulative bytes against a 16 KiB threshold (`FAT_HEADER_THRESHOLD_BYTES = 16 * 1024`). Many boxes hold connection state only within an internal buffer; let that request stream outgrow the buffer mid-flight and the box tears the connection down. The probe clocks the exact byte where it breaks. A reset or timeout after roughly 14 KiB has been sent, or after a response once at least 8 KiB has gone through, is logged as `tcp_16kb_blocked`, because the byte of the break is the signal. It splits resets by timing too: a RST within twice the SYN-ACK round-trip is charged to the in-path node, not the server; one that arrives later is taken as the server itself hanging up, and that gets treated differently. The probe maps each run to one outcome tag: ```text tcp_fat_header_ok request stream reached 16 KiB cleanly tcp_16kb_blocked cut off at the ~14 KiB threshold tcp_freeze_after_threshold stalled past the threshold tcp_reset reset before the threshold tcp_timeout no response tcp_connect_failed never connected tls_handshake_failed TLS setup failed ``` The figure is the decision logic; here is the probe running it, three times against the repo's local-network-fixture, which stands in for the middlebox on loopback: ```text outcome bytesSent rstTimingMs rstOrigin confidence tcp_fat_header_ok 147664 - - none tcp_reset 8273 12 server_rst medium tcp_16kb_blocked 16680 3 server_rst high ``` Three numbers cluster near the threshold and are easy to conflate: 16384 is the 16 KiB cumulative threshold (`FAT_HEADER_THRESHOLD_BYTES`); ~14 KiB is that minus a 2 KiB margin, the point from which the probe reads a teardown as the fat-header signal rather than a plain reset; and 16680 is simply the bytes already sent when this teardown landed, just past the threshold, so `window_cap` fires and it lands as `tcp_16kb_blocked`. On a loopback rig the SYN-ACK RTT is ~0, so the RST falls to `server_rst` either way; only a non-loopback path, with a measurable round-trip, lets the `2×RTT` rule separate `in_path_rst` from `server_rst`. Each probe outcome lands in one of four `ProbeOutcomeBucket` values: `Healthy`, `Attention`, `Failed`, `Inconclusive`. Each carries an event level: `info`, `warn`, or `error`. When the path actively rejected the connection, the failure also carries a `FailureClass`, one of sixteen variants (`DnsTampering`, `TlsAlert`, `HttpBlockpage`, `IpBlockSuspect`, and the rest). `Inconclusive` is the careful bucket. A transient timeout that fired before any real signal arrived goes there and never triggers an automatic strategy change: a switch made on noise would be remembered as the fix. The classification layer collapses all of it into four verdicts, and those decide what happens to the traffic. `TRANSPARENT_OK`: everything works directly, nothing to touch. `OWNED_STACK_ONLY`: the site loads only through the app's own TLS stack, so that's where the connection goes. `NO_DIRECT_SOLUTION`: no on-device packet surgery recovers this one, it needs a tunnel. `IP_BLOCK_SUSPECT`: nothing answers at the IP layer at all. `IP_BLOCK_SUSPECT` is deliberately hard to reach. It needs every DoH-supplied IPv4 address to fail at the SYN, every alternate IPv6 to fail too, and a second independent flow to confirm. Until that confirmation lands, the runner sits in `PendingSecondFlow` and withholds the verdict. A false positive here would shove a connection onto a relay it never needed, so the runner waits for proof. When the verdict does fire, it sets `arm_gate = OwnedStackOnly`: every TLS-family repair is skipped and the engine goes straight to its own stack, the relay path. Rewriting packets is pointless if nothing is home at the address. ## Start with the lightest fix When a verdict calls for packet surgery, a second system takes over. Every fix implements the `DesyncStrategy` trait from `ripdpi-strategy-trait`: `plan` assembles the steps, the other three methods are bookkeeping. The steps are variants of a `DesyncAction` enum, and they share one idea: show the in-path box something other than what the server will see. `Split { offset, disorder }` fragments a TCP segment. `WriteFake { ttl, sni_mode, payload_file }` injects a decoy at a low TTL so it dies in transit before it reaches the server. From there it climbs in weight, from games with the TCP window and TTL up to IP fragmentation and data overlaid on sequence numbers. By default it runs on ordinary unprivileged sockets; the steps that need raw ones live in an optional root helper (`ripdpi-root-helper`) and are skipped quietly when there's no root. A "lightest fix" is concrete. It's a short ordered list of those actions, the output of one strategy's `plan`, applied to a single connection and nothing else. Ten core strategies ship built in, registered through `linkme` distributed slices, so adding one means a single entry and no central `match` statement. The names are utilitarian: `split` fragments a segment, `seq_overlap` lays data over sequence numbers; the rest are in the same vein. Two more, `synack` and `synack_split`, are registered next to them as `Unimplemented` placeholders; SYN-ACK injection runs on a separate path, through the TUN ingress interceptor. The registry tries strategies in registration order and takes the first whose `plan` builds. If applying it fails, an `OnFail` policy decides: roll back to the next one, fall back to plain traffic (the floor when nothing works), or drop the connection. There's a scripted escape hatch too: a feature-gated Lua strategy runs custom logic in a locked-down sandbox (the dangerous stdlib stripped, compiled bytecode refused, a 16 MiB memory ceiling, an instruction-counting watchdog for busy loops, and no escape from its configured directory). Where the action lands inside the flow is not fixed either. A per-flow tuner, `AdaptivePlannerResolver` in `ripdpi-runtime-adaptive`, keeps state per `(network, group, flow kind, target)` tuple and walks five tuning dimensions one at a time when a fix fails (split offset, TLS record offset, and three protocol-specific profiles). The walk order is shuffled per flow from a seed derived from its key, so two flows take different routes instead of stampeding the same path. A win pins the current candidate. A loss benches it for fifteen seconds before it's allowed back in. A layer up, `StrategyEvolver` runs a UCB1 bandit: explore versus exploit. It weights the option that has worked against the one it has sampled least. It scores each strategy combination on success rate, latency, and stability, with a penalty drawn from the failure classes that mean the path actively rejected the connection (`TlsAlert`, `HttpBlockpage`, `Redirect`, `ConnectionFreeze`). Wins decay on a two-hour half-life and losses on a one-hour half-life, so a strategy that worked holds its edge about twice as long as one that failed. A Thompson-sampling alternative sits in the same crate, marked dead code; UCB1 is the default, and the comment says so out loud. Packet surgery is one of two ways out of a bad verdict. The other is `OWNED_STACK_ONLY`: route the connection through the app's own TLS client instead of the system's. That client (`OwnedTlsClientFactory` over the Rust `ripdpi-tls-profiles` crate) keeps verified ClientHello templates for Chrome, Firefox, Safari, and Edge, down to cipher-suite order and session-ticket behavior. It picks one per connection by hashing `SHA-256(authority | session seed | profile set)`, so the choice is stable per host but varies across hosts. It speaks ECH where the target offers it and negotiates the post-quantum hybrid `X25519MLKEM768` group when both ends support it. A checked-in fingerprint snapshot (`owned_stack_tls_fingerprint_snapshot.json`) fails CI if the handshake drifts. ## What the phone remembers The learning is keyed to the network as well as the destination. The `RememberedNetworkPolicyStore`, Kotlin over a Room database, stamps every entry with a SHA-256 hash of the network's scope. The hash folds in transport type, DNS-validation state, captive-portal status, private-DNS mode, the sorted list of DNS servers, and an identity tuple that differs by transport: SSID, BSSID, and gateway for Wi-Fi; operator and SIM codes, carrier ID, and roaming state for cellular. Everything is lowercased and trimmed before it's hashed, and the raw values barely outlive the hash: `CapturedWifiIdentity.toString()` returns `redacted`, the coarse `NetworkSnapshot` used for classification carries no SSID or IP at all, and a repo rule keeps raw SSID and BSSID out of logs and crash reports; the raw SSID never leaves the phone at all. A remembered policy moves through three states: `observed`, `validated`, `suppressed`. Two consecutive failures of a validated policy flip it to suppressed and lock it out for 24 hours; any success resets the failure count and clears the lock. The table keeps at most 64 rows total and forgets anything older than 90 days. On rejoining a known network, the store replays the validated policy at once, then re-checks quietly in the background. Every handoff between Wi-Fi and cellular recomputes the fingerprint and queries the store. Over weeks the phone ends up holding a private map of which networks break which connections, and how. ## Two ways it runs The lighter of the two is proxy mode. `RipDpiProxyService` opens a SOCKS5 proxy on a localhost port; apps that speak SOCKS5 or HTTP CONNECT point at it explicitly, and nothing else on the device is redirected. The other mode runs that same proxy on an ephemeral port, then layers a TUN device on top through Android's `VpnService`. The tunnel reads IP packets off the TUN device (`10.10.10.10/32`, MTU 1500) and opens authenticated SOCKS5 sessions back into the proxy. With no relay configured, the tunnel doesn't change the exit IP: traffic still leaves the device directly. On the way out the packets are only rewritten, so the destination sees the real address and a slightly stranger-looking handshake. When encrypted DNS is on in tunnel mode, an internal FakeIP layer called MapDNS answers queries from the `198.18.0.0/15` range, resolves the real name over the encrypted resolver, and hands the app a synthetic address it pins for the life of the connection. It's never a user-facing toggle: its IPv6 interactions and fail-closed drops make it a poor thing to put behind a switch. That encrypted resolver is its own piece. `ripdpi-dns-resolver` speaks DoH, Oblivious DoH (RFC 9230), and DNSCrypt, so the tool doesn't have to fall back to the system resolver and let whatever the network's DNS answers pollute a measurement. Oblivious DoH splits the knowledge in two: the query goes through a relay that sees the address but not the name, while the target resolver sees the name but not the address, so no single hop sees both. Answers land in a route-aware cache keyed by `(domain, qtype, route decision)`, which forces a fresh lookup when the route changes rather than serving a result resolved for a different path. Routing through your own server is optional, and it comes with caveats. The native core in `libripdpi-relay.so` carries a dozen-odd transports, from Shadowsocks and Trojan up to VLESS Reality and multi-hop chains; WARP and AmneziaWG sit apart, tunnels rather than relays. The list matters less than the line under it in the status doc: every protocol is tested on loopback only, and not one has live remote-endpoint coverage. Mieru shows the gap plainly. It has a native crate and a loopback test, but no activator arm is wired into the profile switch, so a saved profile can't turn it on at this revision. When the path does run through your own server, the configuration travels under a versioned contract. The server-side deploy tooling (`emit-bundle.sh`) emits a standard sing-box JSON config with one extra top-level `ripdpi` object: a `schema_version` plus the fields sing-box has no slot for: an array of AmneziaWG profiles and Hysteria2 obfuscation extras. The app's `SingBoxSubscriptionParser` reads the standard outbounds, then the `ripdpi` block; it rejects a `schema_version` it doesn't know, and a plain sing-box client never notices the key. A contract test on each side holds the shape, `SingBoxRipdpiExtensionParserTest` on the client and the secrets validator on the server, so the format can't drift apart in silence. One credential deliberately never travels this path: a WireGuard private key stays a placeholder (`private_key_placeholder: true`) and is delivered out of band. ## Why the line is drawn where it is The Kotlin/Rust boundary is narrow on purpose. The workspace is 115 crates (the architecture docs still say 114) arranged in nine layers, L0 to L8, and the layering is machine-enforced: a CI script lets only the thirteen top-layer crates touch the `jni` crate or the `android-support` shim. Nothing below them can. Five of those top crates compile to the shared libraries Android loads: `libripdpi.so`, `libripdpi-tunnel.so`, `libripdpi-relay.so`, `libripdpi-warp.so`, `libripdpi-amneziawg.so`. Nothing in the data plane crosses into Java. SOCKS5 sessions, the TUN packet pump, the desync mutations, relay transport, DNS forwarding: all of it stays in Rust. JNI is crossed only to start and stop a session, poll telemetry about once a second, push a network snapshot, and call `VpnService.protect()` on a socket. Every one of those crossings is wrapped. The `android-support` crate's `ffi_boundary` runs each export inside `catch_unwind`, so a Rust panic comes back as a sentinel value instead of unwinding across the `extern "system"` edge into undefined behavior. The JNI build profile even sets `panic = "unwind"` deliberately, because the project's release profile is set to abort, and inheriting that would take down the whole process instead of letting the boundary catch it. Why Rust, specifically, for code that parses untrusted bytes straight off the wire is the subject of the next piece. The same box is still on the path, behaving exactly as before: it buffers the traffic and tears the connection down past the threshold, regardless of which app opened the connection. What changes is the phone. It no longer treats every failure as the same failure: it files a verdict, then tries the smallest intervention that clears the stall and remembers whether it held. The next time that connection hangs on that network, the store already holds the fix that cleared it. --- ### Сброс конкретных соединений при исправной сети: точечная диагностика Published: Jun 2026 | Category: Networking | Tags: DPI, Network Diagnostics, Rust, Android Через мобильный интернет Telegram не поднимает соединение: клиент висит на «Connecting…». Через Wi-Fi тот же клиент подключается мгновенно. Между попытками меняется только сеть: точка подключения и оператор. Этого достаточно, чтобы трафик до одного и того же сервера в одной сети шёл, а в другой обрывался на рукопожатии. Сеть при этом формально исправна: DNS резолвится, другие сайты открываются, ping до 8.8.8.8 ходит без потерь. Отваливаются конкретные соединения, и всегда одни и те же. Это почерк DPI: коробка на стороне оператора разбирает трафик и роняет соединения, которые подходят под сигнатуру. Остальное проходит. В операторских сетях такая асимметрия давно норма. Мобильный оператор не ограничивается маршрутизацией: он снимает отпечатки TLS- и QUIC-рукопожатий, ограничивает скорость отдельных соединений, занижает MTU и сбивает ECN; промежуточная коробка глушит соединение, которое домашний роутер пропустил бы не глядя. В итоге одно направление мертво, соседнее живо, и любой глобальный тумблер «включить везде» ошибётся хотя бы для одного из них. Большинство инструментов в этой нише начинают с догадки. Прогоняют один трюк на уровне пакетов по списку хостов и надеются, что прокатит. Туннели «всё через сервер» уходят в обратную крайность: заворачивают весь трафик на удалённый узел и платят за это задержкой даже там, где ничего не ломалось. И те, и другие назначают лечение, не поставив диагноз. RIPDPI сначала ставит диагноз. ## Диагноз Диагностику ведёт цепочка из четырёх Rust-крейтов: `ripdpi-diagnostics-candidates` готовит входные данные для проб, `ripdpi-diagnostics-probes` задаёт трейт `Probe`, который реализует каждая проверка, `ripdpi-diagnostics-classification` превращает сырые наблюдения в вердикт, а `ripdpi-diagnostics-runner` гоняет всю обойму. Проб больше десятка: целостность и подмена DNS, доступность доменов и QUIC, проверка ECH-рукопожатия, доступность MTProto для Telegram, пропускная способность, опрос DoH-JSON-резолверов. Захардкоженного сервера нет ни у одной: цель передаётся в рантайме через `ProbeContext`, и проба бьёт ровно по тому адресу, который реально пытались открыть. `TcpRunner` открывает одну TLS-сессию и шлёт до 16 padded HTTP HEAD-запросов, каждый крупнее предыдущего, и сверяет накопленный объём с порогом 16 КиБ (`FAT_HEADER_THRESHOLD_BYTES = 16 * 1024`). Многие коробки держат состояние соединения лишь в пределах внутреннего буфера; стоит этим запросам не поместиться в буфер — коробка рвёт соединение. Проба засекает, на каком именно байте всё ломается. Сброс или таймаут после примерно 14 КиБ отправленных данных (или после ответа, когда прокачалось хотя бы 8 КиБ) помечается как `tcp_16kb_blocked`, потому что значим именно байт обрыва. Сбросы она различает и по времени: RST в пределах удвоенного RTT по SYN-ACK списывается на промежуточный узел; если он приходит позже — значит, соединение сбрасывает уже сам сервер, и лечится такое иначе. Проба сводит каждый прогон к одному тегу исхода: ```text tcp_fat_header_ok сессия дошла до 16 KiB без обрыва tcp_16kb_blocked обрыв на пороге ~14 KiB tcp_freeze_after_threshold зависание за порогом tcp_reset сброс до порога tcp_timeout нет ответа tcp_connect_failed не подключился tls_handshake_failed TLS не установился ``` Схема показывает логику решения; а вот она в реальном прогоне — три запуска пробы на локальной сетевой фикстуре репозитория, которая на loopback заменяет промежуточную коробку: ```text outcome bytesSent rstTimingMs rstOrigin confidence tcp_fat_header_ok 147664 - - none tcp_reset 8273 12 server_rst medium tcp_16kb_blocked 16680 3 server_rst high ``` С этим порогом связаны три числа, которые легко перепутать. 16384 — сам порог в 16 КиБ (`FAT_HEADER_THRESHOLD_BYTES`). ~14 КиБ — порог минус 2 КиБ запаса: начиная с этого объёма проба трактует обрыв как признак толстого заголовка. 16680 байт успело уйти к моменту обрыва, чуть за порогом, поэтому срабатывает `window_cap`, и исход — `tcp_16kb_blocked`. На loopback-стенде RTT после SYN-ACK ≈ 0, поэтому любой RST классифицируется как `server_rst`; отделить `in_path_rst` от `server_rst` правилом `2×RTT` можно только на сетевом пути с измеримым RTT. Каждый исход пробы попадает в одну из четырёх категорий `ProbeOutcomeBucket`: `Healthy`, `Attention`, `Failed`, `Inconclusive`. Исходу присваивается и уровень события: `info`, `warn` или `error`. Если соединение на пути активно отбросили, отказу присваивают класс из `FailureClass`, один из шестнадцати (`DnsTampering`, `TlsAlert`, `HttpBlockpage`, `IpBlockSuspect` и прочие). `Inconclusive` — осторожная категория. Случайный таймаут, сработавший до первых осмысленных данных, уходит сюда и не запускает автоматическую смену стратегии: переключение по шуму запомнилось бы как лечение. Слой классификации сводит всё это к четырём вердиктам — они и определяют, что будет с трафиком. `TRANSPARENT_OK`: напрямую всё работает, трогать нечего. `OWNED_STACK_ONLY`: сайт открывается только через собственный TLS-стек приложения — туда соединение и уходит. `NO_DIRECT_SOLUTION`: никакая операция над пакетами на устройстве этот адрес не вытащит, нужен туннель. `IP_BLOCK_SUSPECT`: на уровне IP не отвечает никто. До последнего вердикта добраться намеренно трудно: нужно, чтобы ни один IPv4-адрес из DoH и ни один запасной IPv6 не ответили на SYN, и чтобы это подтвердило второе независимое соединение. Пока подтверждения нет, раннер сидит в `PendingSecondFlow` и вердикт не выносит. Ложное срабатывание здесь загнало бы соединение на ненужный relay, поэтому раннер ждёт доказательств. Когда вердикт всё же выносится, раннер ставит `arm_gate = OwnedStackOnly`: TLS-приёмы пропускаются, и движок сразу переходит к собственному стеку, то есть к relay. Переписывать пакеты бессмысленно, если по адресу никого нет. ## Самое лёгкое средство Когда вердикт требует операции над пакетами, включается вторая система. Каждое средство реализует трейт `DesyncStrategy` из крейта `ripdpi-strategy-trait`: `plan` собирает сами шаги, остальные три метода — служебные. Шаги — варианты enum'а `DesyncAction`, и идея у всех одна: показать промежуточной коробке не то, что увидит сервер. `Split { offset, disorder }` дробит TCP-сегмент. `WriteFake { ttl, sni_mode, payload_file }` подсовывает обманку с заниженным TTL, чтобы та сгорела в пути и до сервера не доехала. Дальше — тяжелее: от игр с TCP-окном и TTL до IP-фрагментации и наложения данных по номерам последовательности. По умолчанию это работает на обычных непривилегированных сокетах; то, что требует сырых сокетов, вынесено в опциональный root-хелпер (`ripdpi-root-helper`) и молча пропускается, когда root недоступен. «Самое лёгкое средство» — вещь конкретная: короткий упорядоченный список таких действий, результат работы `plan` одной стратегии, применённый к одному соединению и больше ни к чему. Десять стратегий ядра встроены и регистрируются через distributed slices из `linkme`, так что добавить новую — одна запись и никакого центрального `match`. Имена утилитарные: `split` дробит сегмент, `seq_overlap` накладывает данные по номерам последовательности; остальные в том же духе. Ещё две, `synack` и `synack_split`, зарегистрированы рядом с ними заглушками `Unimplemented`: вброс SYN-ACK идёт другим путём, через перехватчик на входе TUN. Реестр пробует стратегии по порядку регистрации и берёт первую, у которой `plan` собрал шаги. Если применить не вышло, политика `OnFail` решает: откатиться к следующей, перейти на обычный трафик (последний вариант, если не сработало ничего) или сбросить соединение. Свой сценарий тоже можно написать: Lua-стратегия (под feature-флагом) запускает скрипт в изолированной песочнице (урезанная stdlib, скомпилированный байткод не принимается, лимит памяти 16 МиБ, watchdog по числу инструкций, без выхода за пределы своего каталога). В какой точке соединения сработает действие — тоже не зафиксировано. Tuner для каждого соединения, `AdaptivePlannerResolver` из крейта `ripdpi-runtime-adaptive`, хранит состояние по кортежу `(сеть, группа, тип соединения, цель)` и при неудаче поочерёдно перебирает пять параметров (сдвиг split, сдвиг TLS-записи и три протокол-специфичных профиля). Порядок перебора перемешивается на основе сида, полученного из ключа соединения, так что два соединения идут разными путями. Победа фиксирует текущего кандидата, а поражение откладывает его на пятнадцать секунд, прежде чем снова пустить в дело. Этажом выше работает обучающийся слой. `StrategyEvolver` гоняет многорукого бандита UCB1, классический алгоритм «исследуй или используй», который балансирует между «бери то, что уже работало» и «попробуй то, что пробовал реже всего». Он оценивает каждую комбинацию стратегий по доле успехов, задержке, стабильности и штрафу за детектируемость, а сам штраф вычисляется по тем классам сбоев, что означают «путь активно отверг соединение» (`TlsAlert`, `HttpBlockpage`, `Redirect`, `ConnectionFreeze`). Победы затухают с периодом полураспада в два часа, поражения — в один час, так что сработавшая стратегия держит своё преимущество примерно вдвое дольше провалившейся. В том же крейте лежит альтернатива на Thompson sampling, помеченная как мёртвый код; по умолчанию работает UCB1, и так и написано в комментарии. Операция над пакетами — один из двух путей при плохом вердикте. Второй — `OWNED_STACK_ONLY`: направить соединение через собственный TLS-клиент приложения, а не системный. Этот клиент (`OwnedTlsClientFactory` поверх Rust-крейта `ripdpi-tls-profiles`) хранит выверенные шаблоны ClientHello для Chrome, Firefox, Safari и Edge — вплоть до порядка шифронаборов и поведения session-ticket. Один шаблон на соединение он выбирает по хешу `SHA-256(authority | seed сессии | набор профилей)`: для одного хоста выбор стабилен, между хостами — различается. Он умеет ECH, когда его предлагает целевой сервер, и согласует пост-квантовую гибридную группу `X25519MLKEM768`, когда её поддерживают обе стороны. Зафиксированный снимок фингерпринта (`owned_stack_tls_fingerprint_snapshot.json`) роняет CI, если рукопожатие меняется. ## Что телефон запоминает Результаты этого обучения привязаны и к адресу назначения, и к сети. `RememberedNetworkPolicyStore` (Kotlin поверх базы Room) помечает каждую запись SHA-256-хешем области сети. В хеш входят тип транспорта, состояние валидации DNS, статус captive-portal, режим private DNS, отсортированный список DNS-серверов и кортеж идентичности, зависящий от транспорта: SSID, BSSID и шлюз для Wi-Fi; коды оператора и SIM, carrier ID и состояние роуминга — для сотовой. Перед хешированием всё приводят к нижнему регистру и убирают пробелы, а сырые значения хранятся только до вычисления хеша: `CapturedWifiIdentity.toString()` возвращает `redacted`, обобщённый `NetworkSnapshot` для классификации не несёт ни SSID, ни IP, а правило репозитория не пускает сырые SSID и BSSID в логи и краш-репорты; сырой SSID вообще не покидает телефон. Запомненная политика проходит через три состояния: `observed`, `validated`, `suppressed`. Два провала валидированной политики подряд переводят её в suppressed и запирают на 24 часа; любой успех обнуляет счётчик провалов и снимает блокировку. Всего таблица хранит не больше 64 строк и забывает всё старше 90 дней. Стоит вернуться в знакомую сеть — и хранилище сразу применяет валидированную политику, а потом тихо перепроверяет в фоне. На каждом переходе между Wi-Fi и сотовой сетью отпечаток пересчитывается, и хранилище опрашивается заново. За недели у телефона складывается собственная карта того, какие сети ломают какие соединения и как именно. ## Два режима работы Прокси-режим — тот, что полегче. `RipDpiProxyService` поднимает SOCKS5-прокси на localhost-порту; приложения, которые умеют SOCKS5 или HTTP CONNECT, указывают на него явно, а остальной трафик идёт напрямую. Второй режим запускает тот же прокси на эфемерном порту, а сверху накладывает TUN-устройство через Android VpnService. Туннель читает IP-пакеты с TUN-устройства (`10.10.10.10/32`, MTU 1500) и устанавливает с прокси аутентифицированные SOCKS5-сессии. Без настроенного relay-сервера туннель не меняет внешний IP: трафик всё так же уходит с устройства напрямую. Пакеты на выходе лишь переписываются, так что адресат видит реальный адрес и рукопожатие, которое выглядит чуть необычнее. Когда в туннельном режиме включён шифрованный DNS, внутренний FakeIP-слой под названием MapDNS отвечает на запросы адресами из диапазона `198.18.0.0/15`, резолвит настоящее имя через шифрованный резолвер и отдаёт приложению синтетический адрес, который закрепляет на время соединения. Пользователю это тумблером не показывают: из-за возни с IPv6-режимом и fail-closed-отбрасыванием выносить такое в отдельный переключатель не стали. Этот шифрованный резолвер — отдельная часть: `ripdpi-dns-resolver` умеет DoH, Oblivious DoH (RFC 9230) и DNSCrypt, так что инструменту не нужно откатываться к системному резолверу, и ответы местного DNS не портят измерение. Oblivious DoH делит знание надвое: запрос идёт через relay, который видит адрес, но не имя, а целевой резолвер — имя, но не адрес, так что ни один узел не видит обе половины сразу. Ответы ложатся в route-aware-кэш с ключом `(домен, qtype, решение о маршруте)`. При смене маршрута ключ уже другой: вместо ответа, полученного для другого пути, уходит новый запрос. Маршрут через собственный сервер — опция со своими ограничениями. В нативном ядре `libripdpi-relay.so` — с десяток транспортов, от Shadowsocks и Trojan до VLESS Reality и многоузловых цепочек; WARP и AmneziaWG стоят отдельно, это туннели, а не relay. Важнее списка строчка под ним в статус-документе: каждый протокол проверен только на loopback, живого удалённого эндпоинта нет ни у одного. Mieru — показательный случай. Нативный крейт и loopback-тест на месте, а активатора в переключателе профилей нет, так что из сохранённого профиля Mieru на этой ревизии не включить. Если всё же ходить через свой сервер, конфигурация передаётся по версионированному контракту. Серверная часть деплоя (`emit-bundle.sh`) отдаёт стандартный sing-box JSON с одним дополнительным объектом верхнего уровня `ripdpi`. Туда вынесено всё, для чего в формате sing-box полей нет: `schema_version`, массив профилей AmneziaWG и параметры обфускации Hysteria2. Парсер приложения `SingBoxSubscriptionParser` читает стандартные outbounds, затем блок `ripdpi`; блок с незнакомой версией схемы (`schema_version`) он отвергает, а обычный sing-box-клиент ключ просто не замечает. За контракт отвечают тесты с обеих сторон — `SingBoxRipdpiExtensionParserTest` на клиенте, валидатор секретов на сервере, — так что незаметно рассинхронизироваться сторонам не дадут. Один секрет сознательно не передаётся по этому пути: приватный ключ WireGuard остаётся заглушкой (`private_key_placeholder: true`) и доставляется по отдельному каналу. ## Почему граница проведена именно здесь Границу между Kotlin и Rust намеренно сводят к минимуму. Воркспейс — 115 крейтов (в документах по архитектуре всё ещё 114), разложенных на девять слоёв, от L0 до L8, и слоистость контролируется автоматически: CI-скрипт разрешает трогать крейт `jni` или прослойку `android-support` только тринадцати крейтам верхнего слоя; крейтам ниже это запрещено. Пять из этих верхних крейтов компилируются в разделяемые библиотеки, которые грузит Android: `libripdpi.so`, `libripdpi-tunnel.so`, `libripdpi-relay.so`, `libripdpi-warp.so`, `libripdpi-amneziawg.so`. Вся работа с данными остаётся в Rust и в Java не попадает: SOCKS5-сессии, перекачка пакетов TUN, desync-мутации, relay-транспорт, проброс DNS. Границу JNI пересекают только чтобы запустить и остановить сессию, опросить телеметрию примерно раз в секунду, отдать снимок состояния сети и вызвать `VpnService.protect()` на сокете. И каждое из этих пересечений обёрнуто: функция `ffi_boundary` из крейта `android-support` запускает каждый экспорт внутри `catch_unwind`, так что Rust-паника возвращается как sentinel-значение, а не разматывает стек через границу `extern "system"`, что привело бы к неопределённому поведению. Профиль сборки JNI даже специально ставит `panic = "unwind"`, потому что release-профиль проекта выставлен на abort, и унаследовать это поведение значило бы уронить весь процесс вместо того, чтобы панику перехватил `catch_unwind` на границе. Зачем именно Rust для кода, который разбирает недоверенные байты прямо из сети, — тема следующего текста. Та самая коробка на пути никуда не делась и ведёт себя как прежде: буферизует трафик и рвёт соединение при превышении порога, независимо от того, какое приложение его открыло. Меняется поведение телефона. Он больше не считает все сбои одинаковыми: выносит вердикт, потом пробует минимальное вмешательство, которое снимает затык, и запоминает, помогло ли. В следующий раз, когда это соединение зависнет в той же сети, в хранилище уже записана сработавшая политика. --- ### RAG breaks earlier than people think Published: Apr 2026 | Category: Architecture | Tags: RAG, LLM, Knowledge Management, Architecture ## 1. RAG breaks earlier than people think You can feel the ceiling before you can measure it. At a few hundred thousand documents, a well-tuned vector index starts returning near-misses on queries it answered perfectly at a few thousand. You add a reranker and the top-1 moves back. You add hybrid search and the long tail gets better. Then you keep growing and the failures come back. Same kind, just harder to reproduce. Most teams read this as tuning work. Weller et al. (ICLR 2026) offer a different explanation: a single-vector retriever has a representational ceiling, and past it no amount of downstream cleverness compensates. My own stack hit that ceiling before I could read the explanation for it, and the wiki that replaced it started as a reaction, without a plan. The paper ([arXiv:2508.21038](https://arxiv.org/abs/2508.21038)) states the ceiling as an inequality. For a corpus of $n$ documents with top-$k$ queries and score margin $\gamma$, the embedding dimension $d$ must satisfy $d \geq \frac{\log \binom{n}{k}}{\log(1 + 1/\gamma)}$. Below that bound, some top-$k$ combinations are representationally unreachable in the vector space. The bound works in one direction only: if the embedding dimension sits below the threshold for a given corpus size, no reranker, hybrid search, or prompt engineering recovers the missing combinations. The geometry is already wrong. The failures that follow cluster into three layers. Geometry (the ceiling above) is structural: no tuning moves it. Chunking and context utilisation are budget problems: careful preprocessing helps, but the budgets themselves are shrinking. Attention and hard negatives are generator failures: better prompts help, until the prompts stop mattering. The empirical fit across seven models puts the usable corpus ceiling around 250 million documents at 4096 dimensions, and around 1.7 million at 768, still the common open-weights default. Above those sizes, some top-$k$ sets sit outside the span of any vector the retriever can produce. Most benchmarks live inside those budgets and never probe the edge. LIMIT reaches it at a fraction of those sizes, through combinatorics: each query needs one specific pair of documents back together, so a thousand queries demand a thousand distinct top-2 sets. Fifty thousand documents, sentences like "Jon likes apples". E5-Mistral, GritLM, Qwen3 (the 2025 state of the art) land below 20% Recall@100 at 4096 dimensions on that set. On the same metric BM25 reaches 93.6%, and GTE-ModernColBERT, which keeps a vector per token rather than one per document, reaches 54.8%. The failure sits in the single-vector geometry. A training-time embedding represents only a finite number of distinct top-$k$ sets, and when a query lands outside that set, nothing further in the pipeline compensates. Chunking matters about as much as the model choice. Vectara's 2025 study ([arXiv:2410.13070](https://arxiv.org/abs/2410.13070)) asks whether semantic chunking justifies its cost and finds that in most settings it doesn't beat fixed-size chunking on F1@5 — the retrieval wins on recall, the generator loses on local context, and the two cancel. The failure is the common one: chunks small enough to retrieve precisely are too small for the generator to answer from. Over half the retrieved snippets can be dropped without harming answer quality ([arXiv:2511.17908](https://arxiv.org/abs/2511.17908)). Prior work the paper builds on — RULER and the Context Rot studies — pegs the useful slice of a 128k-token context at 10-20% of what's nominally there. The model skims. Position bias is architectural. The MIT analysis in [arXiv:2502.01951](https://arxiv.org/abs/2502.01951) sketches the mechanism: causal masking accumulates attention on early tokens through every layer, and rotary positional embeddings add a long-term decay to the score. Independent empirical work on long contexts (RULER, Context Rot) confirms the shape in practice — middle-of-context evidence is under-weighted, and longer windows don't buy more attention; the extra tokens land in the zone the model already skips. Hard-negative documents (the near-miss cases a good retriever is designed to surface) degrade end-to-end accuracy; [arXiv:2401.14887](https://arxiv.org/abs/2401.14887) shows the mirror-image result, that random unrelated documents improve accuracy by up to 35%. The U-curve of answer quality against retrieved chunk count is OP-RAG's core finding ([arXiv:2409.01666](https://arxiv.org/abs/2409.01666)). Fixing one failure mode surfaces the next. Contextual Retrieval rewrites chunks with surrounding context before embedding: 35% failure reduction for embeddings alone, 49% when paired with BM25, 67% with a reranker on top. GraphRAG prioritises query-focused summarisation at indexing costs high enough that Microsoft itself shipped LazyGraphRAG as a cheaper follow-up, with comparable quality on global queries at more than 700 times lower query cost. Self-RAG and CRAG move retrieval policy into the model itself. None of them change the basic loop: embed, search, read, generate, repeat. Once per question, paying the full cost again. RAG is the correct first answer to "how do I ground an LLM in my data." For a prototype, a proof of concept, a first integration, it works cheaply and buys you the shape of the problem. It breaks once you try to run a product on it. The embed-search-read-generate loop pays the full retrieval cost on every question and caps out where the geometry does, so scaling the loop scales the bill without moving the ceiling. Two branches. One keeps optimising the query-time loop: better rerankers, late interaction, hybrid search, learned retrieval policies. Most public work is there. The other compiles the corpus once into something the model can read directly. Pay the work upfront, when the source comes in. All of these failures share one assumption: the corpus stays raw, and the model re-derives its understanding of it on every query. Karpathy posted a gist in April 2026 calling this second branch an **LLM Wiki**. ## 2. The wiki reframes the loop > Instead of just retrieving from raw documents at query time, the LLM incrementally builds and maintains a persistent wiki [...] This is the key difference: the wiki is a persistent, compounding artifact. > > — Andrej Karpathy, LLM Wiki gist I read this the first time on a Sunday and reread it twice looking for where the hard part was. The mechanics fit on one page. The wiki has three layers. - **Raw sources** at the bottom: PDFs, transcripts, pasted notes, web clips. Immutable once saved. - **The wiki** in the middle: markdown pages the LLM writes, one per concept or decision, linked with wikilinks. - **The schema** on top: CLAUDE.md, AGENTS.md, whatever file states what the wiki is for. The LLM reads it before every action. **Ingest** does the actual rewriting. A source comes in, the LLM reads it against the schema and the index, and edits whichever pages the source touches, often ten or fifteen at once. Query looks lighter on paper: read the index, drill into a few pages, answer with citations. In practice query also writes back, because an answer worth keeping turns into a new page, and a query can kick off a small secondary ingest to save it. **Lint** is the background pass and the unresolved design question. It walks the wiki for contradictions, stale claims, orphan pages, and missing cross-references. Nobody agrees on when it runs, how aggressive it should be, or what it does when it finds something. Every shipped implementation picks its own answers; mine (frontmatter schema validation at commit time, a linter that only flags values in inline code, append-only logs) are in Part 5. Retrieval is incidental in this model. `index.md`, a plain markdown catalogue maintained by the LLM itself, works at the scale most projects actually live at (around a hundred sources, a few hundred pages) and removes the need for an embedding stack entirely. You can add vectors later if the catalogue stops scaling. You start without them. **Compilation** is the analogy. Each ingest is an incremental build over the previous artefact rather than a fresh pass. A single source typically touches several pages, and a page integrates evidence from several sources. At query time the generator reads compiled pages instead of raw text. Two passages in the gist carry most of the argument: "The human's job is to curate sources, direct the analysis, ask good questions, and think about what it all means. The LLM's job is everything else," and "Humans abandon wikis because the maintenance burden grows faster than the value. LLMs don't get bored, don't forget to update a cross-reference, and can touch 15 files in one pass." The first operation I tried to think through was lint, because I knew the first wiki I'd write would be wrong in ways I wouldn't notice. Karpathy's version is informal: run it periodically, read the report. He cites Vannevar Bush's Memex (1945) as prior art for the whole pattern: a private, curated knowledge store with associative trails between documents, maintained by the person who used it. The original spec ignores time. It treats every piece of content as equally true forever. A lifecycle extension bolts onto the wiki. Facts carry a confidence score that decays over time unless refreshed (roughly the shape of Ebbinghaus's forgetting curve) and resets on access, so claims nobody revisits grow less trusted automatically. When a claim changes, the new version supersedes the old one and carries a pointer back. The old version stays in place. Edges between pages are typed: uses, depends on, contradicts, supersedes. It's shipped at least three times, on very different substrates. ### The shoestring version Bash scripts against an Obsidian vault. A handful of agents, a set of skills, and Bash adapters that compile one source tree into configs for four LLM CLIs; the one hard contract is which tools each agent can touch. Every agent declares its primitives from a fixed set — one might be allowed to read pages and write the index, but forbidden to touch the schema file. Anything off the list fails before the skill ever runs. This design moves almost nothing to ingest time. The agents integrate on demand, and query-time work stays close to what you'd pay without the wiki. What breaks first is cross-platform skew: the same vault driving several LLM CLIs only stays coherent because the contract is narrow enough to hide their differences. ### The heavy version Tauri runtime, React front end, installable binaries. This version packs the whole thing into a desktop application and moves the most work to ingest. A source arrives, the pipeline runs a two-step chain-of-thought pass: analysis against the current purpose file and index, then generation of typed FILE blocks for everything the source touches. Ingest is keyed by content hash. Re-ingest is a no-op. Jobs run through a persistent crash-recoverable queue with retry so an interrupted ingest resumes instead of being lost. Optional vector search bolts onto the catalogue at the end. The ingest cost buys one thing: the wiki is already compiled by the time you open the app. ### The portable version A plugin: one Agent Skills definition, every major LLM CLI, no build step. A single small markdown file — the hot cache — holds it together, a few hundred words of the recent context of work in progress. Session start reads it. Session end updates it. The gap between those two reads is where continuity lives. Without it, every new session opens with amnesia. The cache goes stale if work shifts between contexts without a session-end update. Nothing detects the staleness. ## 3. Other shapes of the same move The three systems in Part 2 all commit to the wiki shape: compile the corpus into markdown, let the human read the artefact. Others solve the same problem from different angles, and what they keep as the artefact says as much as their mechanics. Four families keep coming up. ### Memory agents: Letta and Mem0 Letta (formerly MemGPT, [arXiv:2310.08560](https://arxiv.org/abs/2310.08560)) pages between main context, recall, and archival tiers by function call. Mem0 extracts entity-and-relation facts from every message, resolves conflicts, and writes to a hybrid vector-plus-graph backend; on LOCOMO it reports 66.9% accuracy against OpenAI memory's 52.9% (Mem0 paper, [arXiv:2504.19413](https://arxiv.org/abs/2504.19413)), closing the gap to a full-context baseline to about six points. ### Temporal graphs: Zep Zep does the most interesting thing in this group. Its graph layer, Graphiti ([arXiv:2501.13956](https://arxiv.org/abs/2501.13956)), timestamps every edge with `valid_at` and `invalid_at`. Old beliefs stay in the store with an explicit expiry. A query can ask what the system thought last Tuesday. That's the same problem the wiki's lifecycle extension tries to solve, from the database side. ### Graph RAG: Microsoft's line Graph-RAG starts expensive. Microsoft's GraphRAG ([arXiv:2404.16130](https://arxiv.org/abs/2404.16130)) extracts an entity graph, runs Leiden clustering, and pre-writes community summaries at every level, all with LLM calls before the first query arrives. Microsoft's own follow-up, LazyGraphRAG (Microsoft Research blog, November 2024), drops the pre-summarisation: its indexing costs 0.1% of GraphRAG's, and on global queries it matches GraphRAG's answer quality at more than 700 times lower query cost. LightRAG ([arXiv:2410.05779](https://arxiv.org/abs/2410.05779)) arrives nearby: entity-and-relation extraction at ingest, dual-level retrieval, lightweight per-query profile by design. ### Concept graphs: HippoRAG HippoRAG ([arXiv:2405.14831](https://arxiv.org/abs/2405.14831)) is the strangest. It turns the corpus into a concept graph of noun phrases and answers queries by running Personalized PageRank. Single-step multi-hop reasoning. The v1 failure was entity-centric indexing: the concept graph stripped surrounding context at both ingest and inference, which hurt plain factual recall. HippoRAG 2 ([arXiv:2502.14802](https://arxiv.org/abs/2502.14802)) targets that gap. Letta, Mem0, Zep, and every graph system above store their artefact in a database or index the reader never opens. The wiki is just markdown files. ## 4. When the wiki doesn't fit A wiki isn't always the right answer. Four places it loses: three concrete, one still a suspicion. **Scale.** In my experience, below roughly 50,000 tokens (a soft threshold that depends on which model's window you're paying for) the corpus fits inside a modern context window and the wiki loses to full-context. The wiki starts being worth building around the point where the context stops holding the whole corpus, which is also the point where you have to start maintaining the compression. The upper bound is hundreds of sources and a few hundred pages, above which the plain markdown index stops scaling as a catalogue, and the pattern has to grow its own hierarchy or an embedding layer to keep working. **Error accumulation.** Ingest feeds back into the wiki. Karpathy is explicit about this: partial-context updates miss dependencies, compression drops nuance you can't recover. The feedback works through the ingest input. A pass reads existing wiki pages alongside the new source, so a slightly wrong summary already in the corpus becomes the authority the next ingest integrates against, and the error is inside the corpus by the time lint runs. Lint is the prescribed fix. It catches contradictions and orphan pages cheaply. What it can't see is a quietly wrong summary that nothing downstream notices. Two shipped systems have already corrected for this in public. HippoRAG 2 ([arXiv:2502.14802](https://arxiv.org/abs/2502.14802)) explicitly rewrites v1's entity-centric indexing because it lost context during both ingest and inference. LazyGraphRAG is Microsoft Research conceding, in a product blog, that GraphRAG's up-front indexing costs "may be prohibitive for some users and use cases". **Specific wording and multi-author corpora.** Regulated content that relies on specific wording loses when the wiki paraphrases it. Plain vector RAG over the originals preserves the phrase the wiki dropped. Multi-author coordination breaks things differently. Collaborative Memory ([arXiv:2505.18279](https://arxiv.org/abs/2505.18279)) layers typed read/write permissions over shared memory to keep per-user views isolated, machinery a single-author wiki doesn't need. **Silent rot (suspicion, not diagnosis).** Without a `last_verified` timestamp on each fact, the wiki can't tell which claims still hold and which have quietly gone stale. Zep's bi-temporal edges handle supersession cleanly — a new fact replaces an old one and the edge carries the expiry — but the harder case is the fact nothing in the system actively rechecks. No general answer exists for that. A wiki that stops being maintained doesn't fail loudly. It starts lying, and you find out the next time you read the page. ## 5. What breaks first when you build a wiki The vault behind this part serves an AI infrastructure team: I built its structure and automation, and the team works in it day to day. When you build a wiki-shaped knowledge base it fails in a specific order. I learned the rules for each failure by shipping the wrong fix first. Free-form decision pages don't survive being queried across. Writing decisions without a template feels fine, right up until you ask a question like "what decisions depend on the choice to stay on Python" and you realise `depends_on` is in some pages and not others, sometimes as a list and sometimes as prose. Templates matter for one reason: they turn a pile of pages into a corpus you can query across. A frontmatter schema has to be enforced at commit time, because voluntary schemas decay. Wikilinks break next. You can write perfectly readable prose that names a decision and never links to it. You can do that for weeks. Then you try to walk the graph. The graph is half there: body prose names concepts the frontmatter doesn't, frontmatter names pages the body doesn't, and the link structure is whatever an author happened to remember that morning. Retroactive relinking is cheap in wall time if you have a graph walker and expensive in attention if you don't. The rule I ended up with is that references and links go into the page in the same edit. The hooks enforce only part of it: a link that doesn't resolve fails the commit, but a concept named without a link passes, because nothing checks for missing links. Missing links get added by `/relink`, an agent pass over the pages; its first batch run across the vault inserted 110 links in 45 files in one commit. Writing links in the same edit keeps that pass small. My first instinct with decision pages was the ADR tradition: write once, supersede with a new document when the underlying call changes. That works when decisions are rare and the record is legal. In a live wiki it turns the corpus into a thicket. You end up with two files, and a reader has to know which one is live. The better shape is one file per topic, current state at the top, and the chronology moved off the page into an append-only log that records every change as a timestamped line, plus whatever version control already gives you. One piece of history stays on the page: the options a decision turned down, with the reason, because a reader needs them to judge the call. This is uncomfortable for anyone who learned decision-writing from compliance culture. It works anyway. Some rules are right only with the right scope. My inbox-policing hook runs at commit time and checks only what the commit stages: a raw transcript can wait unprocessed in the staging folder for a whole session, and a commit fails only if it tries to add that file without a `type:` classification. The rule is sound, because staging should not turn into a graveyard, but a hook that checked the whole folder would block every commit until the transcript was processed. The same rule also exists as a merge-gate CI job, written but commented out until the local hook has shown it works reliably. Commit-time, push-time, and merge-time aren't the same tool. The first schema linter I wrote was too trusting with prose. It scanned every token in page bodies and flagged anything that looked like an enum value from the frontmatter schema. It caught invented values correctly. It also flagged the word "active" in a sentence, and "draft" in the phrase "first draft", and blocked commits that were completely fine. The rewrite only scans inline code spans. Values wrapped in backticks are machine-readable, and prose is left alone. Making the linter narrow is what kept it from being turned off. Single-agent assumptions break early. An agent rulebook that grew up against one platform's permission model, tool surface, and skill invocation style won't port when a second agent arrives. Shimming one agent's interface onto the other's rulebook papers over a structural change that has already happened. The shape that survives is two peer rulebooks and a hook that refuses commits when the mirrors drift. Every skill gets written against both rulebooks in the same edit. Cost per skill is higher. The payoff: either agent can pick up the vault cold. That matters more than expected once the work has to outlast a particular CLI's session. During the hardening phase, metadata iterates faster than content. The rule files, the schema documents, the audit log: those are the most-edited artefacts. At first this looks like a smell. Governance churning more than the corpus it governs is counter-intuitive. Then it stops looking like a smell. Content accumulates quietly. The rules around content evolve fast, because real requirements surface only once real content exists. Metadata churn therefore tracks the vault's state: while the rules keep changing, the vault is still being fitted to real content, and when they stop, it has either stabilised or fallen out of use. The uncomfortable part of this approach is that it sounds like overhead. The maintenance cost is paid upfront in templates, schema, hook chains, and the shape of the append-only log, and the per-commit cost afterwards is close to zero. A long meeting becomes a dozen pre-linked decision pages in the time it takes to run the ingest pass, because the graph walkers and the schema do the work the author would otherwise do by hand. Not every edit, though. Irreversible ones still require a timestamped log entry that names me. The `decided_by` field in a decision's frontmatter is always a human name; the agents never fill it. Contested claims stay in place with counter-evidence added below them, never silently overwritten. I don't have a good way to tell whether the wiki is improving. Add-rate is easy to measure and doesn't mean what I'd like it to mean. The public benchmarks ask a different question than I want answered: LongMemEval ([arXiv:2410.10813](https://arxiv.org/abs/2410.10813)) covers five memory abilities across long chat histories, LOCOMO ([arXiv:2402.17753](https://arxiv.org/abs/2402.17753)) tests multi-session conversational memory, and DMR from MemGPT ([arXiv:2310.08560](https://arxiv.org/abs/2310.08560)) scores multi-session recall. None of them asks the question I care about — whether the same query asked six months apart on the same evolving corpus returns a consistent answer — and Zep's paper reporting on DMR ([arXiv:2501.13956](https://arxiv.org/abs/2501.13956)) admits the questions are ambiguous enough that a high score can reflect LLM inference skill rather than memory fidelity. No published benchmark measures time-travel consistency: the same question asked at different points as the corpus evolves. Summarisation-fidelity metrics like FaithEval exist. No memory system paper reports using them. What the commit hooks catch is mechanical: schema violations, broken wikilinks, contradictions flagged by the narrative-schema linter. Silent regression, a quietly wrong summary that nothing in the system notices, shows up the next time I open the page. Or it doesn't. ## 6. What to build, and when The choice between RAG and a wiki is a choice about where to pay. RAG stands up fast, costs little per token, and caps out when the corpus gets large enough that the retrieval geometry runs out of room. A wiki costs more at ingest and less at query, holds up better under repeated questions on the same material, and rewards the human who maintains it. Real systems tend to ship both layers: vector search over immutable sources for precise recall, a compiled wiki for synthesis across a specific project. What the failure modes in this piece share is where compute gets paid. The move from query time to ingest time is the underlying bet. Whether the output lands as markdown pages, typed edges in a temporal graph, or pre-computed community summaries depends on what the artefact is for. A wiki is read by humans; a graph is queried by systems. Both sit downstream of the same decision: compile once, read many. I don't have a neat ending. The wiki I built will rot in places I stop re-reading, and the commit hooks will keep catching schema violations while silent summaries drift. That's the trade the pattern makes: no durability guarantee, just a failure mode that lives in a file a person can open. The measure I'd want — same query, same corpus, six months apart, do the answers agree — isn't something any benchmark reports on yet, so for now the signal is the next time I open a page and flinch at what it says. --- ### RAG ломается раньше, чем кажется Published: Apr 2026 | Category: Architecture | Tags: RAG, LLM, Knowledge Management, Architecture ## 1. RAG ломается раньше, чем кажется Потолок начинаешь чувствовать раньше, чем его получается измерить. На нескольких сотнях тысяч документов хорошо настроенный векторный индекс начинает промахиваться на запросах, на которые отлично отвечал при нескольких тысячах. Добавляешь реранкер — top-1 возвращается. Добавляешь гибридный поиск — длинный хвост выравнивается. Растёшь дальше — и отказы возвращаются, те же по сути, только воспроизвести их сложнее. Большинство команд считает это задачей тюнинга. Weller и соавторы (ICLR 2026) объясняют иначе: у одновекторного ретривера есть потолок выразимости, и выше него никакие ухищрения в конце пайплайна уже не помогают. Мой собственный стек упёрся в этот потолок раньше, чем я узнал, чем он объясняется, и вики, которая заменила этот стек, появилась как реакция, без плана. В статье Weller и соавторов потолок записан как неравенство. Для корпуса из $n$ документов, top-$k$ запросов и зазора оценок $\gamma$ размерность эмбеддинга $d$ должна удовлетворять условию $d \geq \frac{\log \binom{n}{k}}{\log(1 + 1/\gamma)}$. Ниже этой границы часть top-$k$ комбинаций просто нельзя представить в векторном пространстве. Граница работает в одну сторону: если размерность эмбеддинга ниже порога для этого размера корпуса, ни реранкер, ни гибридный поиск, ни промпт-инжиниринг не вытащат пропущенные комбинации. Геометрия уже неправильная. Дальнейшие режимы сбоя ложатся в три слоя. Геометрия (описанный выше потолок): структура, её тюнингом не сдвинуть. Чанкинг и утилизация контекста: вопрос бюджета — аккуратная предобработка лечит, но сами бюджеты усыхают. Внимание и hard-negatives: провалы на стороне генератора, лучшими промптами лечатся, пока промпты вообще работают. По эмпирической аппроксимации на семи моделях потолок полезного корпуса — около 250 миллионов документов при размерности 4096 и около 1,7 миллиона при 768 (до сих пор стандарт для открытых моделей). Выше этих объёмов часть top-$k$-комбинаций уже не может представить ни один вектор, который строит энкодер. Большинство бенчмарков на эту границу не выходят, а LIMIT выходит на неё при куда меньшем корпусе, за счёт комбинаторики: каждому запросу нужна своя пара документов, и всего таких пар тысяча. В нём пятьдесят тысяч документов и предложения вроде «Jon likes apples». Лучшие эмбеддинг-модели падают ниже 20% Recall@100 при размерности 4096. BM25 по той же метрике выдаёт 93,6%. Проблема в геометрии одного вектора: обученный эмбеддинг представляет конечное число top-$k$-наборов, и когда требуемая запросом комбинация в их число не входит, ничего дальше по цепочке ситуацию не исправляет. Следующим ломается чанкинг, и весит он столько же, сколько выбор модели. Исследование Vectara ([arXiv:2410.13070](https://arxiv.org/abs/2410.13070)) ставит вопрос, окупается ли семантический чанкинг, и на F1@5 отвечает: в большинстве сценариев не окупается — ретривер выигрывает на recall, генератор проигрывает на локальном контексте, и два эффекта гасят друг друга. Провал обычный: чанки достаточно мелкие для точного поиска оказываются слишком мелкими, чтобы по ним ответить. Аттеншн обычно недооценивают. Больше половины сниппетов можно выкинуть без ущерба для ответа ([arXiv:2511.17908](https://arxiv.org/abs/2511.17908)). Работы, на которые эта статья опирается — RULER и исследования Context Rot — оценивают полезный срез 128k-токенового контекста в 10–20% от номинала. Модель скользит по поверхности. Смещение по позиции зашито в архитектуру. Разбор MIT ([arXiv:2502.01951](https://arxiv.org/abs/2502.01951)) даёт механизм: каузальная маска копит внимание на первых токенах, а rotary embeddings добавляют долгосрочное затухание в score. Независимая эмпирика на длинных контекстах (RULER, Context Rot) подтверждает эту форму на практике — середина контекста недоучитывается, и длинное окно не прибавляет внимания; лишние токены попадают в зону, которую модель и так пропускает. Ошибки складываются. Hard-negative документы (похожие на нужные, но нерелевантные; хороший ретривер как раз их и находит) ухудшают итоговую точность. А случайные несвязанные документы, как показывает [arXiv:2401.14887](https://arxiv.org/abs/2401.14887), её улучшают — максимум на 35%. Починка одной проблемы обычно вскрывает следующую. Улучшения есть, и все настоящие. Contextual Retrieval переписывает чанки с учётом окружающего контекста перед эмбеддингом: –35% провалов на одних эмбеддингах, –49% с BM25, –67% с реранкером сверху. GraphRAG делает упор на саммари сообществ и платит за это индексацией, настолько дорогой, что Microsoft сам выпустил LazyGraphRAG как дешёвое продолжение: на глобальных запросах качество сопоставимое, а запрос обходится более чем в 700 раз дешевле. Self-RAG и CRAG переносят политику извлечения внутрь самой модели. Ни одно из этих улучшений не меняет базовый цикл: эмбеддинг, поиск, чтение, генерация. Каждый запрос проходит цикл заново, платя полную цену. RAG — правильный первый ответ на вопрос «как прицепить LLM к своим данным». Для прототипа, PoC, первой интеграции он работает, стоит недорого и даёт форму задачи. Ломается, когда на нём пытаются держать продукт. Цикл эмбеддинг-поиск-чтение-генерация платит полную стоимость извлечения на каждый вопрос и упирается в тот же потолок, где упирается геометрия; масштабирование цикла масштабирует счёт, а не потолок. Отсюда две ветки. Одна продолжает улучшать цикл в момент запроса: реранкеры, late interaction (ColBERT и наследники), гибридный поиск, обучаемые политики извлечения. Ветка продуктивна, и публичной работы там больше. Другая ветка устроена иначе. Вместо того чтобы платить за извлечение, чанкинг и внимание на каждом вопросе, корпус один раз компилируется во что-то, чем модель может пользоваться сразу. Вся работа делается заранее, в момент поступления источника. Все эти режимы стоят на одном допущении: корпус остаётся сырым, и модель каждый раз заново выводит, что с ним делать. Karpathy в апреле 2026 года опубликовал гист, где назвал эту вторую ветку **LLM Wiki**. ## 2. Вики переносит работу на другой конец цикла > Вместо того чтобы при каждом запросе искать по сырым документам, LLM инкрементально строит и поддерживает вики, которая сохраняется между запросами [...] В этом главное отличие: вики — долговременный артефакт, ценность которого накапливается. > > — Andrej Karpathy, гист LLM Wiki Я прочитал этот гист в воскресенье и перечитал дважды, пытаясь понять, где подвох. Записка короткая. Вся механика помещается на одну страницу. У вики три слоя. - **Сырые источники** внизу: PDF, расшифровки, заметки, клиппинги с веба. После сохранения не меняются. - **Сама вики** в середине: markdown-страницы, которые LLM пишет, по одной на концепцию или решение, связанные вики-ссылками. - **Схема** сверху: CLAUDE.md, AGENTS.md, любой файл с описанием того, зачем эта вики. LLM читает его перед каждым действием. Основную работу делает **ингест**. Приходит источник, LLM читает его вместе со схемой и индексом и правит страницы, которых источник касается, часто десять-пятнадцать за раз. Формально запрос проще: прочитать индекс, нырнуть в несколько страниц, ответить с цитатами. На деле запрос тоже пишет обратно: если ответ стоит того, чтобы его сохранить, он становится новой страницей, и запрос запускает маленький вторичный ингест. Поиск тут вторичен. Karpathy пишет, что большинству проектов (около сотни источников, несколько сотен страниц) хватает `index.md`, простого markdown-каталога, который LLM ведёт сама. Векторы можно добавить, когда индекса перестанет хватать, а начинать без них. Гист строится на аналогии с **компиляцией**. Источники на входе, вики на выходе, каждый ингест — инкрементальная сборка поверх предыдущего артефакта. Один источник обычно затрагивает несколько страниц. Через несколько ингестов каждая страница вбирает данные из разных источников, и результат перестаёт быть похожим на пересказ отдельного документа. Генератор читает уже скомпилированные страницы, до сырого текста дело не доходит. Почему эту сборку стоит поручить модели, гист объясняет двумя фрагментами: «Задача человека — отбирать источники, направлять анализ, задавать хорошие вопросы и думать, что всё это значит. Задача LLM — всё остальное» и «Люди забрасывают вики, потому что затраты на поддержку растут быстрее, чем польза от неё. LLM не устают от рутины, не забывают обновить перекрёстную ссылку и могут за один проход поправить 15 файлов». Первая операция, которую я попытался продумать, был **линт** — обход вики в поисках противоречий, устаревших утверждений, сирот и сломанных ссылок. Я знал, что первая написанная мной вики окажется неправильной в местах, которые я сам не замечу. У Karpathy линт описан неформально: запускай периодически, читай отчёт. На деле интересный вопрос в том, когда его запускать и что он делает, когда что-то находит. Гист оставляет это на читателя. Каждая реальная реализация выбирает свои ответы; мои (валидация фронтматтер-схемы на коммите, линтер, флагующий только значения в бэктиках, append-only-лог) — в Части 5. Karpathy ссылается на «Мемекс» Вэнивара Буша (1945) как на предшественника. «Мемекс» был задуман как личное курируемое хранилище знаний с ассоциативными связями между документами. Одна система, один человек, свои источники. Исходная спецификация игнорирует время. Она считает любой контент одинаково верным навсегда, а реальные знания так не работают. Очевидное расширение — прикрутить жизненный цикл. Факты получают оценку уверенности, и она убывает без обращений: утверждение, к которому никто не возвращается, теряет доверие само. При замене старая версия остаётся с обратным указателем. Рёбра между страницами типизированы (зависит от, противоречит, замещает), и по ним можно строить структурные запросы. Можно узнать, что вики считала верным в любой момент в прошлом. Файл схемы тут главный. Без него LLM не знает, как себя вести. Всё это ничего не стоит, пока идея существует только в гисте. Она уже запущена как минимум три раза, на разных платформах, и интересно не то, что они все работают, а то, куда каждая перекладывает работу. ### Бюджетная версия Система построена на Bash-скриптах поверх хранилища Obsidian. В ней горстка агентов, набор навыков и Bash-адаптеры, которые из одного исходного дерева собирают конфигурации для четырёх LLM CLI; жёсткий контракт один — какие инструменты каждый агент может использовать. Каждый агент объявляет свои примитивы из фиксированного набора — одному, например, разрешено читать страницы и писать индекс, но запрещено трогать файл схемы. Объявит что-то за пределами списка — сборка упадёт. Эта конструкция почти ничего не перекладывает на ингест: агенты собирают по запросу, и работы примерно столько же, сколько без вики вообще. Слабое место, на мой взгляд, — сами платформы: одно хранилище на несколько LLM CLI держится только потому, что контракт достаточно узкий и скрывает различия между их интерфейсами. ### Тяжёлый подход Упаковывает всё в десктопное приложение. Tauri-рантайм с React-фронтендом, установочные бинарники для трёх основных платформ. Сюда на ингест перекладывается больше всего работы. Приходит источник, и конвейер ингеста прогоняет двухшаговый chain-of-thought: сначала аналитическое чтение источника на фоне текущего файла целей и индекса, затем генерация типизированных FILE-блоков для всего, что источник затрагивает. Ингест привязан к хешу содержимого, так что повторный ингест ничего не делает, а задания проходят через персистентную очередь с перезапуском. Опциональный векторный поиск прикручен в конце. Весь этот ингест нужен ради одного: вики уже скомпилирована к моменту, когда ты открываешь приложение. Архитектура заточена под один сценарий: очередь не должна терять состояние при сбое. ### Портативная версия Плагин. Одно Agent Skills определение, которое запускается на любом крупном LLM CLI без этапа сборки. Всё держится на одном маленьком markdown-файле — горячем кэше в несколько сотен слов, хранящем свежий контекст текущей работы. В начале сессии он читается, в конце обновляется, и непрерывность держится именно на этом промежутке. Без него каждая новая сессия начинается с амнезии. Кэш протухает, если работа переключается между контекстами без обновления в конце сессии, и ничто в системе этого не замечает. ## 3. Другие формы того же манёвра Три системы из Части 2 делают один шаг — компилируют корпус в markdown и отдают артефакт человеку. Остальные решают ту же задачу с другой стороны, и то, что они сохраняют как артефакт, говорит столько же, сколько их механика. Четыре семейства стоит назвать. ### Агентные системы памяти: Letta и Mem0 Letta (бывший MemGPT, [arXiv:2310.08560](https://arxiv.org/abs/2310.08560)) переключает данные между основным контекстом, recall-хранилищем и архивом через function call. Mem0 извлекает факты entity-and-relation из каждого сообщения, разрешает конфликты и пишет в гибридное хранилище (вектор + граф); на LOCOMO он показывает 66,9% точности против 52,9% у памяти OpenAI (статья Mem0, [arXiv:2504.19413](https://arxiv.org/abs/2504.19413)), сокращая разрыв до full-context-бейзлайна примерно до шести пунктов. ### Временны́е графы: Zep Zep делает самое интересное в этой группе. Его графовый слой, Graphiti ([arXiv:2501.13956](https://arxiv.org/abs/2501.13956)), ставит временны́е метки `valid_at` и `invalid_at` на каждое ребро. Старые убеждения остаются в хранилище с явным сроком годности. Запрос может спросить, что система считала верным в прошлый вторник. Тот же вопрос, на который отвечает расширение жизненного цикла в вики, только со стороны базы данных. ### Граф-RAG: линия Microsoft Граф-RAG-системы стартуют дорого. GraphRAG от Microsoft ([arXiv:2404.16130](https://arxiv.org/abs/2404.16130)) извлекает граф сущностей, запускает Leiden-кластеризацию и заранее пишет саммари сообществ на каждом уровне, причём всё это вызовами LLM ещё до первого запроса. Собственное продолжение Microsoft, LazyGraphRAG (блог Microsoft Research, ноябрь 2024), выкидывает предварительную суммаризацию: индексация обходится в 0,1% стоимости GraphRAG, а на глобальных запросах качество ответов сопоставимое при стоимости запроса более чем в 700 раз ниже. LightRAG ([arXiv:2410.05779](https://arxiv.org/abs/2410.05779)) рядом: извлечение сущностей и отношений на ингесте, двухуровневый поиск, небольшой расход на один запрос по замыслу. ### Концептные графы: HippoRAG HippoRAG ([arXiv:2405.14831](https://arxiv.org/abs/2405.14831)) — самый странный из всех. Он превращает корпус в концептный граф из именных групп и отвечает на запросы, прогоняя Personalized PageRank. Multi-hop-рассуждение за один шаг. Провал v1 — entity-centric-индексация: концептный граф срезал окружающий контекст и на ингесте, и на инференсе, что било по обычному факт-поиску. HippoRAG 2 ([arXiv:2502.14802](https://arxiv.org/abs/2502.14802)) целится именно в этот провал. Вики отличается тем, что результат компиляции — обычный markdown. Файлы, которые человек может открыть, отредактировать и прочитать. Системы памяти прячут артефакт в базу данных, граф-системы — в кластерное дерево или PageRank-скор. Когда что-то идёт не так, вики ломается у тебя на виду. Остальные ломаются за стеной абстракции, и ты узнаёшь об этом по качеству ответов, а не по состоянию хранилища. ## 4. Когда вики не подходит Вики не всегда правильный ответ. Четыре места, где она проигрывает: три конкретных, одно пока подозрение. **Масштаб.** По моему опыту, ниже примерно 50 000 токенов (граница мягкая и зависит от того, чьё окно контекста оплачиваешь) корпус целиком помещается в контекстное окно, и вики проигрывает полному контексту. Запускать компрессию раньше — значит платить за то, что модель и так может обработать целиком. Верхняя граница тоже описана в гисте: Karpathy оценивает рабочий диапазон в сотни источников и несколько сотен страниц, выше чего плоский markdown-индекс перестаёт работать как каталог и паттерну приходится отращивать иерархию или эмбеддинг-слой. **Накопление ошибок.** Karpathy называет этот режим в оригинальном гисте: ингест возвращает результаты обратно в вики, обновления на частичном контексте упускают зависимости, а компрессия теряет нюансы, и восстановить их нельзя. Петля замыкается через вход ингеста. Проход читает существующие страницы вики вместе с новым источником, поэтому слегка ошибочный саммари, уже попавший в корпус, становится авторитетом для следующего ингеста, и ошибка оказывается внутри корпуса к моменту, когда линт добирается до неё. Самое неприятное: такая ошибка выглядит как нормальный текст. Линт тут штатное лекарство. Он ловит противоречия и сирот-страницы дёшево, но тихо ошибочный саммари ему не виден, если ничто ниже по потоку не замечает расхождения. Это уже ломалось публично. HippoRAG переписал свою индексацию во второй версии, потому что первая теряла контекст и при ингесте, и при инференсе. LazyGraphRAG — признание Microsoft Research в продуктовом блоге, что расходы GraphRAG на предварительную индексацию «для некоторых пользователей и сценариев могут оказаться непомерными». **Точные формулировки и несколько авторов.** Регулируемый контент, завязанный на конкретные формулировки, проигрывает, когда вики их перефразирует, — а обычный векторный RAG по оригиналам сохраняет ту фразу, которую вики потеряла. Многоавторская координация давит с другой стороны: Collaborative Memory ([arXiv:2505.18279](https://arxiv.org/abs/2505.18279)) накручивает над общей памятью типизированные read/write-разрешения, чтобы у каждого пользователя был свой срез — механика, без которой одноавторская вики спокойно обходится. **Тихое гниение (подозрение, не диагноз).** Без временно́й метки `last_verified` на каждом факте вики не может определить, какие из её утверждений ещё верны, а какие тихо устарели. Битемпоральные рёбра Zep чисто закрывают замещение — новый факт заменяет старый, ребро несёт срок годности — но тяжёлый случай другой: факт, который ничто в системе активно не перепроверяет. Для него общего ответа нет. Вики, за которой перестали следить, не падает с грохотом. Она начинает врать, и понимаешь это, когда в следующий раз открываешь страницу. ## 5. Что ломается первым, когда строишь вики Хранилище, о котором идёт речь в этой части, обслуживает команду AI-инфраструктуры: структуру и автоматизацию строил я, а команда пользуется им каждый день. Вики-хранилище ломается в определённой последовательности. Правила для каждой поломки я узнал, сначала выкатив неправильный фикс. Первая поломка: свободноформатные страницы решений не выживают, когда по ним начинают делать запросы. Писать решения без шаблона удобно ровно до момента, когда задаёшь вопрос типа «какие решения зависят от выбора остаться на Python» и обнаруживаешь, что `depends_on` есть в одних страницах, а в других нет, где-то списком, где-то прозой. Шаблоны нужны ровно для одного: они превращают кучу страниц в корпус, по которому можно строить запросы. Фронтматтер-схема должна проверяться на коммите, потому что добровольные схемы деградируют. Вики-ссылки ломаются следующими. Можно неделями писать читаемые тексты, в которых решение упоминается по имени, но ни разу не линкуется. Потом пытаешься обойти граф. Граф готов наполовину: тело текста называет концепции, которых нет во фронтматтере, фронтматтер перечисляет страницы без обратных ссылок в теле, а структура ссылок определяется тем, что автор вспомнил в то утро. Перелинковать задним числом недолго, если есть обходчик графа. Без него дорого по вниманию. Правило, к которому я пришёл: ссылки и упоминания пишутся в одной правке. Хуки проверяют только часть этого правила: ссылка, которая никуда не ведёт, валит коммит, а концепт, упомянутый без ссылки, проходит, потому что упоминания без ссылки ничто не ищет. Недостающие ссылки расставляет `/relink`, агентный проход по страницам; первый пакетный запуск по всему хранилищу вставил 110 ссылок в 45 файлов одним коммитом. Когда ссылки пишутся в той же правке, этот проход остаётся маленьким. Мой первый инстинкт с решениями: ADR-традиция. Написал один раз, при изменении создаёшь новый документ, замещающий старый. Работает, когда решения редки, а запись имеет юридическую силу. В живой вики получается чаща. Два файла, и читатель должен знать, какой из них актуален. Лучше один файл на тему, текущее состояние наверху, а хроника вынесена со страницы в append-only лог, где каждое изменение записывается строкой с таймстэмпом, плюс то, что и так хранит система контроля версий. Одна часть истории остаётся на странице: отвергнутые варианты с причинами, потому что без них читатель не оценит само решение. Для тех, кто учился писать решения в культуре комплаенса, это неудобно. Работает тем не менее. Некоторым правилам нужны точные границы. Мой inbox-хук срабатывает при коммите и проверяет только добавляемые в него файлы: сырая расшифровка может лежать в staging-папке необработанной всю сессию, и коммит отклоняется, лишь если в него попадает такой файл без классификации в поле `type:`. Правило здравое, staging не должен превращаться в кладбище, но хук, проверяющий всю папку, блокировал бы любую фиксацию, пока расшифровка ждёт обработки. То же правило есть и в виде CI-джоба на мёрдж-гейте: он написан, но закомментирован, пока локальный хук не покажет, что работает надёжно. Коммит, пуш и мёрдж — разные инструменты. Первый схема-линтер, который я написал, слишком доверял прозе. Он сканировал каждый токен в теле страницы и помечал всё, что выглядело как значение enum из фронтматтер-схемы. Придуманные значения он ловил правильно. Но он также ловил слово «active» посреди предложения и «draft» во фразе «первый черновик», и блокировал коммиты, с которыми всё было в порядке. Переписанная версия сканирует только код в бэктиках. Значения в бэктиках машиночитаемы, проза остаётся нетронутой. Линтер выжил потому, что стал узким. Допущения под одного агента ломаются рано. Свод правил, выросший на модели прав, наборе инструментов и стиле вызова навыков одной платформы, не переносится, когда появляется второй агент. Натягивание интерфейса одного на свод правил другого маскирует структурные изменения, которые уже произошли. Выжившая форма: два равноправных свода правил и хук, отклоняющий коммиты при расхождении зеркал. Каждый навык пишется под оба свода в одной правке. Цена на навык выше. Зато любой из агентов может подхватить хранилище с нуля и работать с ним, и это оказывается важнее, чем ожидалось, когда работа должна пережить конкретную CLI-сессию. На этапе закалки хранилища метаданные итерируются быстрее контента. Файлы правил, документы схемы, лог аудита: это самые редактируемые артефакты, а отдельные страницы решений нет. Поначалу это выглядит как антипаттерн — управление, меняющееся быстрее управляемого корпуса, противоречит интуиции, — а потом перестаёт. Контент накапливается тихо. Правила вокруг контента эволюционируют быстро, потому что реальные требования проявляются только тогда, когда реальный контент уже есть. Поэтому по частоте правок метаданных видно состояние хранилища: пока правила меняются, его ещё подгоняют под реальный контент, а когда перестают меняться, оно либо устоялось, либо заброшено. Хорошего способа определить, становится ли вики лучше, у меня нет. Скорость добавления легко измерить, но она говорит не о том, о чём хотелось бы. Публичные бенчмарки задают другой вопрос: LongMemEval ([arXiv:2410.10813](https://arxiv.org/abs/2410.10813)) покрывает пять способностей памяти на длинных чатах, LOCOMO ([arXiv:2402.17753](https://arxiv.org/abs/2402.17753)) тестирует multi-session-память в диалогах, DMR из MemGPT ([arXiv:2310.08560](https://arxiv.org/abs/2310.08560)) оценивает multi-session-recall. Ни один не спрашивает того, что меня интересует — возвращает ли один и тот же запрос через полгода на том же эволюционирующем корпусе согласованный ответ — и статья Zep по DMR ([arXiv:2501.13956](https://arxiv.org/abs/2501.13956)) признаёт, что вопросы там достаточно неоднозначны, чтобы высокий счёт отражал навыки инференса LLM, а не точность памяти. Ни один не измеряет консистентность во времени — один и тот же вопрос, заданный в разные моменты эволюции корпуса. Метрики достоверности суммаризации существуют, и ни один отчёт по системам памяти их не использует. Коммит-хуки ловят механическое: нарушения схемы, битые ссылки, противоречия. Тихая регрессия обнаруживается в следующий раз, когда открываешь страницу. Или не обнаруживается. Неудобная часть этого подхода в том, что он звучит как оверхед. Стоимость поддержания заплачена заранее в шаблонах, схеме, цепочках хуков и форме append-only лога, а стоимость каждого коммита после этого близка к нулю. Длинная встреча превращается в дюжину пре-линкованных страниц решений за время одного ингест-прохода, потому что обходчики графа и схема делают ту работу, которую автор делал бы руками. Снаружи не видно, что всё в хранилище уже прошло через валидацию к моменту, когда читатель до него добирается. Но это верно не для каждой правки. Необратимые правки по-прежнему требуют записи в логе с моим именем. Поле `decided_by` во фронтматтере решения всегда заполняется человеком; агенты его не трогают. Спорные утверждения остаются на месте, контраргументы добавляются ниже, молчаливая перезапись исключена. ## 6. Что строить и когда Выбор между RAG и вики — выбор, где платить. RAG быстро разворачивается, стоит копейки на запрос и упирается в стену, когда корпус вырастает настолько, что геометрия извлечения перестаёт помещаться. Вики требует больше на ингесте и меньше на запросе, лучше держит повторяющиеся вопросы к тому же материалу и вознаграждает того, кто её ведёт. Реальные системы обычно держат оба слоя: векторный поиск по неизменяемым источникам для точного recall, скомпилированную вики для синтеза по конкретному проекту. Общее у режимов сбоя из этой статьи — где платят за вычисления. Перенос с момента запроса на момент ингеста и есть та самая ставка. Во что превратится артефакт — markdown-страницы, типизированные рёбра темпорального графа, заранее посчитанные саммари сообществ — зависит от того, для чего он нужен. Вики читает человек; граф запрашивает система. Оба ниже по течению от одного решения: компилировать один раз, читать много. Аккуратного финала у меня нет. Вики, которую я построил, частично сгниёт в местах, куда я перестану возвращаться, а хуки продолжат ловить нарушения схемы, пока тихие саммари будут дрейфовать. Вот компромисс, на который идёт паттерн: никакой гарантии долговечности, только форма сбоя, которая лежит в файле, открываемом человеком. Метрика, которую хотелось бы иметь — один и тот же запрос, тот же корпус, через полгода, совпадают ли ответы — ни в одном бенчмарке пока не измеряется, так что до тех пор сигнал прежний: следующий раз, когда я открою страницу и поморщусь от того, что там написано. --- ## Projects ### Copilot AI Platform An AI assistant for retail investors built on a layered multi-agent architecture. Temporal handles durable orchestration and human-in-the-loop checkpoints, LangGraph manages supervisory agent graphs, and PydanticAI provides type-safe agent definitions. Self-hosted vLLM on GKE (H200 GPU) handles sensitive client data, routed through Bifrost AI Gateway to external APIs. Deterministic context layer (MCP servers in Go and Python) separates financial calculations and compliance logic from LLM reasoning. Langfuse provides full observability: traces, evaluations, and cost analytics via OpenTelemetry. Platforms: Backend, AI | Tags: LangGraph, PydanticAI, Temporal, vLLM, Bifrost, Langfuse, FastAPI, Go | Year: 2026 | Status: Active ### CI/CD Slack Bot A Python/Flask/Slack Bolt service built in 6 weeks and evolved over 3 years into the mobile team's primary CI/CD interface. Exposes build triggers, release branch management, version bumping, and pipeline status directly in Slack. Extended with Google Play API v3 integration (AAB publish, track promotion, staged rollout, production monitoring) and Firebase Crashlytics integration via an HTTPS JSON-RPC MCP client. Runs on Kubernetes as a single-replica Socket Mode Bolt service with OpenTelemetry distributed tracing spanning Slack through GitLab CI to Gradle builds. Platforms: Backend | Tags: Python, Flask, Slack Bolt, GitLab CI, Google Play API, OpenTelemetry, Kubernetes | Year: 2023 | Status: Maintained ### AGENTS.md Framework A policy specification framework that gives AI coding agents consistent project constraints without per-session configuration. AGENTS.md files define architecture contracts, coding standards, and never-do rules at workspace and project level. A deterministic rule-routing table maps file globs to dedicated rule documents, ensuring agents load the correct domain constraints before modifying build logic, CI pipelines, or feature code. Includes 24 project-specific executable skills with YAML trigger phrases for automatic selection, and a 4-agent task-proof-loop workflow (spec-freezer, builder, verifier, fixer) with JSON schema-validated evidence artifacts. Platforms: Tooling | Tags: Claude Code, Gemini CLI, Codex, YAML, MCP, AI Agents | Year: 2026 | Status: Active ### Heimdall A Rust toolchain that reads transcripts from Claude Code, Codex, Cursor, OpenCode, Pi, Copilot, Xcode CodingAssistant, Cowork, and Amp, then surfaces an interactive dashboard with cost estimates, 5-hour billing-block burn-rate projection, cache efficiency, task categorization, waste-detection grade, context-window tracking, and rate-limit tracking — all running entirely on the user's machine. Three surfaces ship together: a `heimdall` CLI with embedded web dashboard, a `heimdall-hook` sub-second real-time ingest writing per-tool cost on every PreToolUse (~50 ms p99), and an MCP server exposing 9 analytics tools over stdio + HTTP. A native macOS menu-bar app (Swift) bundles the CLI with browser-session import. Storage is local SQLite; no telemetry, no network egress. Platforms: Tooling, macOS | Tags: Rust, Swift, TypeScript, MCP, SQLite, FinOps, Observability | Year: 2026 | Status: Active ### Kotlin CI Toolchain A set of Kotlin CLI tools built with picocli and Apache SSHD/JGit that replaced all Ruby/Fastlane automation in the Android CI pipeline over 8 weeks. ReleaseManager handles release branch creation, version bump orchestration, and back-merge safety checks. SlackUploader manages artifact upload and Slack notifications. A lint-codequality converter produces GitLab Code Quality reports from Android Lint and Detekt output. The migration unlocked Gradle Configuration Cache across all CI jobs and eliminated all Ruby environment issues. Platforms: Android | Tags: Kotlin, picocli, GitLab CI, Gradle, CLI | Year: 2025 | Status: Maintained ### ANR Watchdog A production Android library for detecting Application Not Responding events at two layers. The Java-level monitor watches for main thread stalls, while a C++ native signal handler (built with NDK/CMake) catches SIGABRT-class events that bypass the JVM entirely. Integrates with ApplicationExitInfo for post-mortem diagnostics and uploads structured reports to Firebase Crashlytics. Feature-flag-gated for controlled production rollout. Requires understanding the JNI boundary, NDK build configuration, ProGuard symbol retention for native binaries, and threading constraints when calling JVM from signal context. Platforms: Android | Tags: Kotlin, C++, NDK, CMake, Firebase Crashlytics | Year: 2024 | Status: Stable ### po4yka.dev This site. Built with Astro 7 and React 19 islands architecture, deployed on Cloudflare Workers with D1 (SQLite) for the database. Admin panel uses passkey-first authentication via WebAuthn. Content pipeline is single-source: MDX for blog posts, JSON for projects and experience, with build-time generators producing static TypeScript data files and seed SQL. Designed with anti-AI-slop principles: minimal, credible aesthetic where typography carries the design. Platforms: Web | Tags: Astro, React, TypeScript, Cloudflare Workers, D1, WebAuthn | Year: 2025 | Status: Active ### RIPDPI An Android application that optimizes network connectivity by routing traffic through a local SOCKS5 proxy with adaptive DPI evasion capabilities. Operates in proxy mode or local VPN redirection mode. Implements adaptive evasion strategies with semantic markers and adaptive split placement. Provides encrypted DNS via DoH/DoT/DNSCrypt, integrated diagnostics with automatic probing and strategy auditing, per-network policy memory with intelligent network handover handling. Native Rust modules handle performance-critical paths via JNI bridges. Includes a user guide generator that creates annotated PDFs from live screenshots. Platforms: Android | Tags: Kotlin, Rust, JNI, SOCKS5, DNS, VPN, NDK | Year: 2025 | Status: Active ### Ratatoskr A self-hosted content reader and summarizer with three first-party clients. The backend (Python, FastAPI) scrapes articles and extracts YouTube transcripts via a multi-provider chain (Scrapling, Firecrawl, Playwright), then generates structured summaries via LLM through OpenRouter. The mobile/desktop client is built with Kotlin Multiplatform + Compose Multiplatform, sharing ~80-90% of business logic across Android, iOS, and Desktop with offline-first delta sync. A React/TypeScript web client doubles as a Telegram Mini App for on-demand summarization from the browser. Features include offline reading, semantic search, channel digests, and MCP server integration for AI agents. Ktor for networking, SQLDelight for persistence, Decompose for navigation, Koin for DI. Platforms: Android, iOS, Backend, Web | Tags: KMP, Compose Multiplatform, Python, FastAPI, Telegram Bot, React, Ktor, SQLDelight | Year: 2025 | Status: Active ## Experience ### AI Engineer -- Garage IT (Apr 2026 — Present) Building Copilot, a multi-agent assistant for a regulated multi-asset investment platform. Own the agent architecture, self-hosted LLM infrastructure, and AI platform decisions. - Designing a layered multi-agent system with Temporal, LangGraph, and PydanticAI, capped at 3-4 agents per cluster based on error-propagation research - Separated deterministic logic (MCP servers in Go/Python) from LLM reasoning - Selected Bifrost AI Gateway over APISIX and Kong for multi-LLM routing - Deployed Langfuse for LLM observability and set up Claude Code for the engineering team ### Senior Android Developer -- Garage IT (Dec 2024 — Present) Decomposed a 1,500-LOC Activity monolith into a plugin architecture, replaced the Ruby CI pipeline with Kotlin tooling, and introduced AI coding agent workflows. - Decomposed the main Activity into 15 standalone lifecycle plugins, stabilized 10+ crash regressions - Unified navigation across 3 app flavors, moved 80+ screens into the shared codebase - Replaced Fastlane/Ruby CI pipeline with Kotlin CLI tools over 8 weeks, unlocking Gradle Configuration Cache - Designed AGENTS.md policy framework across 3 repos for Claude Code, Gemini CLI, and Codex ### Android Developer -- Garage IT (Nov 2022 — Dec 2024) Android and MobileOps engineer on a multi-asset retail investment platform across 3 regulated markets. Owned the CI/CD infrastructure, release pipeline, and internal developer tooling. - Modularized a monolithic codebase into 15+ Gradle modules over 2 years - Led RxJava to Kotlin Coroutines migration across the data layer, shipped behind a Firebase Remote Config toggle - Built the team's CI/CD Slack bot in 6 weeks (Python, Flask, Slack Bolt) ### Android Developer -- VK (Feb 2022 — Nov 2022) Middle Android Developer on VK Clips, a short-form video product inside the VK super-app. - Led the transition from a monolithic codebase to a modularized architecture, splitting VK Clips into separate SDKs - Migrated to Dependency Injection with a custom wrapper over Dagger 2 - Optimized performance using Android Profiler: 2% faster response time, 11% reduction in memory usage ### Junior Android Developer -- VK (Mar 2021 — Feb 2022) Junior Android Developer on VK Clips, a short-form video service at VK. - Accelerated key UI elements by 30%, contributing to a 20% increase in user retention - Built a privacy and user data control system for video content - Developed a curated video collections interface, leading to a 7% increase in user engagement ### Industrial Practice (Internship) -- EPAM Systems (Feb 2021 — Jun 2021) Android development internship focused on Agile practices, automated testing, and cross-platform development. - Built a companion Android app for D&D: character stats tracking, reference library, virtual dice-rolling - Worked in a cross-functional team of 4 using Agile methodologies ### Junior Android Developer (Part-time) -- LETI (Sep 2020 — Mar 2021) Built a Kotlin-based Android app with MVVM architecture as a university course hub. - Used Jetpack Compose (Alpha), LiveData, and Data Binding; trained five team members on Jetpack libraries - Integrated with Moodle via RESTful APIs using Kotlin Coroutines - Achieved 50+ active students within one month; 75% positive feedback ## Skills - **Languages**: Kotlin, Python, Go, Rust, TypeScript - **Android**: Jetpack Compose, Coroutines/Flow, Koin, OkHttp/Retrofit, Gradle, KSP, R8 - **AI/ML**: LangGraph, PydanticAI, Temporal, vLLM, Langfuse, MCP, Claude Code - **Backend**: Flask, FastAPI, Slack Bolt, Google Play API, Firebase - **DevOps**: GitLab CI, Docker, Kubernetes, Helm, OpenTelemetry, GKE - **Tools**: Android Studio, Obsidian, Git, Macrobenchmark, Baseline Profiles