Haystack Europe 2026: the bleeding edge of retrieval engineering
Haystack Europe took place in Berlin on 15 & 16 September as part of Berlin Search Week: two days, single track, about 120 people who build search for a living. About half of the audience said they run an LLM in production!
The sessions broadly shared four underlying themes:
- Agents are a first-class search consumer. Retrieval stacks are being rebuilt to serve two very different query distributions: human users and agents writing long, precise, operator-rich queries.
- Hybrid search is the default. Netflix, S&P Global, Leroy Merlin and Coveo all described dense plus sparse retrieval fused with reciprocal rank fusion (RRF) in production.
- Evaluation is key, and difficult. A lot of the agenda was dedicated to the difficulty in proving that (& how) a retrieval system works.
- New embedding models. Two novel embedding models were presented - one to embed hierarchy and another that has deterministic dimensions to enable query-time composability.
I've written up my notes from the conference below (unfortunately I couldn't attend all of them). All talk recordings are in the Haystack Europe 2026 playlist and the full schedule is on the Haystack website.
Agents as search consumersPermalink to this heading
The unreasonable effectiveness of BM25 for agentic searchPermalink to this heading
Jo Kristian Bergum, Hornet.dev

I was nodding along throughout this talk. Jo's first point was that BM25's weak reputation in benchmarks mostly comes from untuned default parameters, not the algorithm. Also that agents are better searchers than people: they know entity names, dates and jargon, and they happily write long, operator-rich queries which are perfect for lexical matching. He also showed agents working over retrieved documents as a workspace or file system, pulling in detail progressively rather than stuffing everything into context.
On BrowseComp-Plus, GPT-5 with a BM25 retriever scores 55.9% accuracy; swap in Qwen3-Embedding-8B and it reaches 70.1% with fewer search calls. Jo showed that this benchmark is flawed - and that a tuned BM25 algorithm can achieve 95.7% answer-presence. Better retrieval is always relevant, and fewer calls is also a cost and latency win. So basically - test it on your systems, don't be afraid to change the defaults, and choose the approach that best balances accuracy and performance (like all search implementations!).
I've added this to my (growing) list of experiments to run - if you give an agent access to lexical and dense search plus weighted RRF, which gives the best answer in the fewest turns? I avoided the temptation to try to do this in the evening before my talk the next day!
The RAG Cost Curve: When an index beats live searchPermalink to this heading
Event page · Slides · Recording
Simon Hearne, Zilliz - that's me (full write-up)

Mine was the only talk that put a price on the index versus live-search decision, which surprised me given how much of the agenda was about agents. The argument: "RAG or agentic search" is a false binary. It's a cost model, decided by corpus size, update frequency and query volume.
I grounded it in two open-source systems, memsearch (agent memory over Markdown) and claude-context (semantic code search), and in benchmarks of what compression really costs across scalar, product and RaBitQ quantisation.
During the benchmarks I found a few interesting things:
- Agents don't need indexes for code that models have been trained on (e.g. long-running public GitHub repositories)
- For private code, the break-even point for an index is in the tens to hundreds of daily queries
- If the index is ~free (local, serverless, on an existing box) then the break-even point is on the first query!
Agent memory that works with your existing search stackPermalink to this heading
Event page · Slides · Recording
Paul-Louis Nech, Algolia

Algolia added memory to its Agent Studio under one constraint: customers run keyword search, neural search or both, and nobody should need vector-only infrastructure to use it. Paul-Louis showed how far classic search features go for agent memory: tie-breaking ranking, numeric filters, optional word boosting and LLM-generated tags that are boosted at query time. He also covered compressing conversations into multi-query payloads, keyword extraction across 35 languages, and LongMemEval results, including why benchmarks are mostly useless. Again - benchmark on your own data and use real customer feedback.
Paul-Louis' thesis is different to mine: we agree that agent memory is a retrieval problem, not a transcript-replay problem. Where our talks differed is on what the index should hold. Keyword memory is a measurable improvement on no memory for answer-presence, and on context stuffing for latency & token budgets. One of the most interesting moments in the talk (for me) was that removing stop-words degraded answer-presence - it makes sense on reflection: English stop-words include my and their, not and is, few and most. Stripping those words can totally change the meaning of a sentence.
This was the second talk (after Jo's) that made me want to evaluate BM25 vs. vector search for memory infrastructure. Both claude-context and memsearch use hybrid search so I need to separate the queries to determine answer-presence for each individually compared to the default RRF.
Hybrid search as the defaultPermalink to this heading
From Tries to Transformers: Scaling Semantic SearchPermalink to this heading
Event page · Slides (PDF) · Recording
Ivan Provalov, Netflix

Ivan shared some great statistics from Netflix that were new (and surprising) to me: only 16% of Netflix discovery starts with search, and the average query is three characters long (typing on TVs is hard). At that length there are no semantics to embed: a user typing wes could be looking for Wes Anderson, West Side Story or Westerns.
Netflix's answer is to stop picking one interpretation. They generate every plausible facet mapping, retrieve for each, and merge the lists with reciprocal rank fusion (plus personalisation), with a "did you mean" helper on top. The stack has moved from trie lookups to a Joint-BERT intent model and a hybrid lexical plus embedding engine, running at 25k QPS (!). Ivan also shared Netflix's open source query testing framework Netflix/q, which he demoed in a lightning talk.
Short-query ambiguity is a query understanding problem (something I looked at for typeahead search). Fusion and ranking based on inferred intent is a good solution to a complex UX problem. I'd love to learn how Netflix use customer feedback (CTR, re-search rate etc.) and metrics like MRR to measure and refine their search system.
Query Understanding with Late Interaction & Wormhole VectorsPermalink to this heading

Trey's opening point was that most teams lean on ranking maths and hope it compensates for never really understanding the query. The talk was then a toolkit for fixing that upstream.
The headline idea is wormhole vectors. Standard hybrid search runs a lexical query and a dense query in parallel, then fuses the lists. A wormhole instead travels between spaces: run the lexical query, pool the dense vectors of the documents it matched (synthesising a new, pool-averaged vector), and search the dense space from that point. It also works in reverse, from dense results back into sparse terms via a semantic knowledge graph. Around it he covered late interaction models and collaborative filtering for learning latent features from user behaviour.
Trey reports benchmarks beating conventional hybrid search (Wormhole precision@10 of 0.862 vs RRF precision@10 of 0.795). If that holds up broadly, it changes what a hybrid search API should expose: not just two queries and a fusion weight, but result-set pooling as a first-class operation. Another thing to add to your benchmarks and evaluations!
As both BM25 and dense search have got cheaper and faster, this concept is realistic in production with low double digit millisecond latency budgets. I would like to test it vs RRF with oversampling and semantic re-ranking.
Sparse encoders: bridging products and knowledge graphsPermalink to this heading
Event page · Slides (PDF) · Recording
François Gaillard, ADEO (Leroy Merlin, Bricoman)

ADEO runs home-improvement stores in Europe such as Leroy Merlin. ADEO's products have complex hierarchy and concepts (building materials vs. tools etc), so they need every product and query mapped onto its knowledge graph of concepts. Millions of product-concept pairs were poorly aligned, and a meaningful share of queries (18M / 21%) did not map correctly.
Dense embeddings didn't help: off-the-shelf models don't encode the same concepts, and the results couldn't be explained when a business rule misfired. This lack of 'explainability' is an issue I have encountered many times before, especially in e-commerce and legal document search scenarios.
François presented their solution to this challenge: a custom SPLADE-style sparse encoder which injects knowledge-graph concepts straight into the model vocabulary as tokens. Every sparse vector is then expressed in the catalogue's own vocabulary. A generative T5 expansion approach lost on inference time (600 ms against 60 ms), but latency was reduced by adding FLOPS-loss and tuning the weights to push sparsity past 99.9%, without losing semantic depth.
This is the most practical case for learned sparse retrieval I've seen. You get semantic matching with concept understanding, a vector you can actually read and debug (it's just a list of terms), and it can be supported by any search engine that can index sparse vectors (such as Milvus).
SPGE Enhanced Search: relevancy engine for Agentic RAGPermalink to this heading
Event page · Slides (PDF) · Recording
Tomasz Florin & Artur Barczyński, S&P Global

S&P Global runs one hybrid relevancy platform for two very different consumers: the search box in its core enterprise product, and the retrieval layer behind agentic RAG workloads. Lexical and semantic retrieval are fused with RRF, and business-driven relevance rules shape the final ranking.
Tomasz and Artur did a good job of calling out the trade-offs in deploying a single retrieval stack. Complex business data needs precision and explainability as much as recall, and agents and analysts want different things from the same index. It's a clear picture of where large, regulated enterprises are heading: one retrieval platform, several consumers, with relevance rules the business can still control.
New embedding modelsPermalink to this heading
Hyperbolic embedding models for visual searchPermalink to this heading
Event page · Recording · Paper

Most image search stores hierarchy outside the vectors, in tags or a knowledge graph. hyper³labs trained an open image-text model (Hyper3-CLIP) in hyperbolic space instead, where distance from the origin encodes specificity: "animal" sits near the centre, "border collie" near the edge. The hierarchy lives in the geometry, so a single nearest-neighbour query can respect it.
The weights are open, and the results show good gains over the industry-standard CLIP multi-modal embedding model. The overall concept is intriguing and I can see many applications for it, from the obvious e-commerce category search to more interesting use cases like security incidents or fraud cases - where the specificity dimension could encode the maturity or depth of the attack / case.
I was so impressed with the concept that I have built it out on Milvus - watch this space for the writeup!
Introducing per-document embedding fine-tuningPermalink to this heading
Philippe Bouzaglou, Vectra

Embeddings are usually frozen at index time. They hold whatever the model could infer from the input, and changing them means fine-tuning the model and re-embedding the whole collection (which isn't cheap!). Vectra's e-commerce model is novel in that it has deterministic dimensions for attributes such as colour or material, so you can fine-tune an individual document's embedding at query time, without having to tune the model and re-generate the embeddings.
The demo that stuck with me: a product embedded from its photo, then enriched with facts the photo can't show, like organic cotton, a hippie look and festival wear. The same mechanism lets you tune embeddings in real time to a user's preferences.
If composable embeddings catch on, vectors stop being a static snapshot of a document and become something you update. That changes how we should think about index updates, and it's a very different workload from periodic bulk re-embedding.
Evaluating search systemsPermalink to this heading
Beyond LLM Judges: Deep Evaluation for Conversational SearchPermalink to this heading
Event page · Slides (PDF) · Recording
Ruchi Juneja, MediaMarktSaturn

MediaMarktSaturn's conversational shopping assistant had just gone into an A/B test, and Ruchi walked through why scoring only the final answer with an LLM judge wasn't enough. Their framework evaluates each layer separately: did the agent extract the right search terms, did its facets match the catalogue, how often did it hit zero results.
The standout finding was that implicit filters hurt. For example, if a customer mentioned "256GB", the agent turned it into a hard filter and narrowed the results, even though that shopper might have happily bought a 128GB or 512GB model. It's a small case of a big problem for agentic retrieval: an agent that follows the query too literally over-constrains it.
Evaluation is still largely human-graded on recall, accuracy and reliability, with business metrics not yet wired in. Her closing line is worth stealing: a good evaluation doesn't just tell you how good the system is, it tells you what to fix next.
Ruchi also described how they distilled their evaluation into a "Conversational Search Reliability Score" graded from Poor (< 50) to Excellent (≥ 90) so they could track progress and report to leadership. I love a simple metric like this, but I'm also curious whether the metric correlates with business metrics (conversion rate, abandonment etc.), which wasn't shared in the talk.
Teaching your shopping assistant to learn from its mistakesPermalink to this heading
Event page · Slides (PDF) · Recording

This was a great companion to Ruchi's talk. When OTTO launched its product discovery assistant, failures were silent: shoppers just gave up. Jens described the weekly loop they now run over tens of thousands of production sessions:
- Collect sessions and merge in the funnel events the customer actually saw.
- Judge each one with an LLM anchored to those events.
- Cluster failures into buckets with a single root cause each.
- Review clusters against real dialogues to weed out LLM hallucinations.
- Act on retrieval, prompt engineering or guardrails, then confirm in the next run.
The failures OTTO surfaced were interesting (and actionable): over-constrained queries returning zero results, gaps in catalogue attributes as well as purple team activity trying to break the guardrails. The key takeaway was to use LLMs to generate hypotheses and initial classification to improve the initial results, but use real customer interactions as well as human review to fine-tune in production.
What I'm taking awayPermalink to this heading
The centre of gravity in search has moved up the stack. Getting vectors into an index is assumed; the hard problems are understanding the query, serving agents and humans from one retrieval layer, and proving that the answer was right. The financial cost of all of this wasn't discussed at all outside my talk - but nothing is free. I feel like token costs and compute might become a limiting factor as these solutions scale.
I also have a lot of experiments and benchmarks to run!