# Haystack Europe 2026: the bleeding edge of retrieval engineering

> Agents as search consumers, hybrid retrieval as the default, and how to evaluate search systems - a summary of my takeaways from Haystack Europe 2026.

Author: Simon Hearne (https://simonhearne.com/about/)
Published: 2026-09-29
Canonical: https://simonhearne.com/2026/haystack-europe-2026/
Tags: search, VectorDB, Milvus, conference

---
[Haystack Europe](https://haystackconf.com/) took place in Berlin on 15 & 16 September as part of [Berlin Search Week](https://berlinsearchweek.com/): two days, single track, about 120 people who build search for a living. About half of the audience said they run an LLM in production!

The sessions broadly shared four underlying themes:

- **Agents are a first-class search consumer.** Retrieval stacks are being rebuilt to serve two very different query distributions: human users and agents writing long, precise, operator-rich queries.
- **Hybrid search is the default.** Netflix, S&P Global, Leroy Merlin and Coveo all described dense plus sparse retrieval fused with reciprocal rank fusion (RRF) in production.
- **Evaluation is key, and difficult.** A lot of the agenda was dedicated to the difficulty in proving that (& how) a retrieval system works.
- **New embedding models.** Two novel embedding models were presented - one to embed hierarchy and another that has deterministic dimensions to enable query-time composability.

I've written up my notes from the conference below (unfortunately I couldn't attend all of them). All talk recordings are in the [Haystack Europe 2026 playlist](https://www.youtube.com/playlist?list=PLM1qgHLmgOk0) and the full schedule is on the [Haystack website](https://haystackconf.com/schedule/).

## Agents as search consumers

### The unreasonable effectiveness of BM25 for agentic search

[Event page](https://haystackconf.com/session/the-unreasonable-effectiveness-of-bm25-for-agentic-search/) · [Recording](https://www.youtube.com/watch?v=ZnmtJKFonLk)

*[Jo Kristian Bergum](https://www.linkedin.com/in/jo-bergum/), [Hornet.dev](https://hornet.dev/)*

![bar chart comparing query length: human queries median 2, p99 8 terms; agent queries median 10, p90 17 terms](https://simonhearne.com/images/haystack-europe-2026/bm25_agents_hornet.png)
*The median agent query is longer than 99% of human queries (Hornet)*

I was nodding along throughout this talk. Jo's first point was that BM25's weak reputation in benchmarks mostly comes from untuned default parameters, not the algorithm. Also that agents are better searchers than people: they know entity names, dates and jargon, and they happily write long, operator-rich queries which are perfect for lexical matching. He also showed agents working over retrieved documents as a workspace or file system, pulling in detail progressively rather than stuffing everything into context.

On [BrowseComp-Plus](https://arxiv.org/abs/2508.06600), GPT-5 with a BM25 retriever scores 55.9% accuracy; swap in Qwen3-Embedding-8B and it reaches 70.1% with fewer search calls. Jo showed that this benchmark is flawed - and that a tuned BM25 algorithm can achieve 95.7% answer-presence. Better retrieval is always relevant, and fewer calls is also a cost and latency win. So basically - test it on your systems, don't be afraid to change the defaults, and choose the approach that best balances accuracy and performance (like all search implementations!).

I've added this to my (growing) list of experiments to run - if you give an agent access to lexical and dense search plus weighted RRF, which gives the best answer in the fewest turns? I avoided the temptation to try to do this in the evening before my talk the next day!

### The RAG Cost Curve: When an index beats live search

[Event page](https://haystackconf.com/session/the-rag-cost-curve-when-an-index-beats-live-search/) · [Slides](https://talks.simonhearne.com/rag-cost-curve/) · [Recording](https://www.youtube.com/watch?v=Ipqnh4upvrI)

*[Simon Hearne](https://www.linkedin.com/in/simonhearne/), Zilliz - that's me ([full write-up](/2026/rag-cost-curve/))*

![chart of cost per correct answer for no retrieval, agentic and indexed retrieval on known and unseen code](https://simonhearne.com/images/haystack-europe-2026/rag_cost_curve.png)
*Cost per correct answer: indexes win on code the model has never seen*

Mine was the only talk that put a price on the index versus live-search decision, which surprised me given how much of the agenda was about agents. The argument: "RAG or agentic search" is a false binary. It's a cost model, decided by corpus size, update frequency and query volume.

I grounded it in two open-source systems, [memsearch](https://github.com/zilliztech/memsearch) (agent memory over Markdown) and [claude-context](https://github.com/zilliztech/claude-context) (semantic code search), and in benchmarks of what compression really costs across scalar, product and [RaBitQ](https://arxiv.org/abs/2405.12497) quantisation.

During the benchmarks I found a few interesting things:

- Agents don't need indexes for code that models have been trained on (e.g. long-running public GitHub repositories)
- For private code, the break-even point for an index is in the tens to hundreds of daily queries
- If the index is ~free (local, serverless, on an existing box) then the break-even point is on the first query!

### Agent memory that works with your existing search stack

[Event page](https://haystackconf.com/session/agent-memory-that-works-with-your-existing-search-stack/) · [Slides](https://plnech.github.io/talks/haystack26/talk.html#1) · [Recording](https://www.youtube.com/watch?v=3aKXXRrCkk8)

*[Paul-Louis Nech](https://www.linkedin.com/in/plnech/), Algolia*

![table comparing four memory strategies by seconds, input tokens and facts recalled; memory as a tool uses 9,184 tokens to recall one fact](https://simonhearne.com/images/haystack-europe-2026/agent_memory.png)
*Four approaches to agent memory: memory as a tool uses 27× the tokens yet gives less context (Algolia)*

Algolia added memory to its [Agent Studio](https://www.algolia.com/products/ai/agent-studio) under one constraint: customers run keyword search, neural search or both, and nobody should need vector-only infrastructure to use it. Paul-Louis showed how far classic search features go for agent memory: tie-breaking ranking, numeric filters, optional word boosting and LLM-generated tags that are boosted at query time. He also covered compressing conversations into multi-query payloads, keyword extraction across 35 languages, and [LongMemEval](https://arxiv.org/abs/2410.10813) results, including why benchmarks are mostly useless. Again - benchmark on your own data and use real customer feedback.

Paul-Louis' thesis is different to mine: we agree that agent memory is a retrieval problem, not a transcript-replay problem. Where our talks differed is on what the index should hold. Keyword memory is a measurable improvement on no memory for answer-presence, and on context stuffing for latency & token budgets. One of the most interesting moments in the talk (for me) was that removing stop-words degraded answer-presence - it makes sense on reflection: English stop-words include `my` and `their`, `not` and `is`, `few` and `most`. Stripping those words can totally change the meaning of a sentence.

This was the second talk (after [Jo's](#the-unreasonable-effectiveness-of-bm25-for-agentic-search)) that made me want to evaluate BM25 vs. vector search for memory infrastructure. Both claude-context and memsearch use hybrid search so I need to separate the queries to determine answer-presence for each individually compared to the default RRF.

## Hybrid search as the default

### From Tries to Transformers: Scaling Semantic Search

[Event page](https://haystackconf.com/session/from-tries-to-transformers-scaling-semantic-search/) · [Slides (PDF)](https://pretalx.com/media/haystackeu26/submissions/7F3XA8/resources/From_Tries_to_gXhR9ny.pdf) · [Recording](https://www.youtube.com/watch?v=dW2vP_Jl7MM)

*[Ivan Provalov](https://www.linkedin.com/in/provalov/), Netflix*

![diagram of a five-step pipeline: query logs, human annotations, LLM-as-judge, quality gate, train and deploy](https://simonhearne.com/images/haystack-europe-2026/netflix.png)
*The data engine behind Netflix's query understanding model*

Ivan shared some great statistics from Netflix that were new (and surprising) to me: only 16% of Netflix discovery starts with search, and the average query is three characters long (typing on TVs is hard). At that length there are no semantics to embed: a user typing `wes` could be looking for Wes Anderson, West Side Story or Westerns.

Netflix's answer is to stop picking one interpretation. They generate every plausible facet mapping, retrieve for each, and merge the lists with reciprocal rank fusion (plus personalisation), with a "did you mean" helper on top. The stack has moved from trie lookups to a [Joint-BERT](https://arxiv.org/abs/1902.10909) intent model and a hybrid lexical plus embedding engine, running at 25k QPS (!). Ivan also shared Netflix's open source query testing framework [Netflix/q](https://github.com/Netflix/q), which he demoed in a [lightning talk](https://www.youtube.com/watch?v=HJJfMubvEcE).

Short-query ambiguity is a query understanding problem (something I looked at for [typeahead search](/2026/fast-typeahead-milvus/)). Fusion and ranking based on inferred intent is a good solution to a complex UX problem. I'd love to learn how Netflix use customer feedback (CTR, re-search rate etc.) and metrics like MRR to measure and refine their search system.

### Query Understanding with Late Interaction & Wormhole Vectors

[Event page](https://haystackconf.com/session/query-understanding-with-late-interaction-wormhole-vectors/) · [Recording](https://www.youtube.com/watch?v=XeA27uhpY_4)

*[Trey Grainger](https://www.linkedin.com/in/treygrainger/), [Searchkernel](https://searchkernel.com/)*

![slide showing wormhole vectors connecting sparse lexical, dense semantic and behavioral vector spaces](https://simonhearne.com/images/haystack-europe-2026/wormhole_vectors.png)
*Wormhole vectors traverse between sparse, dense and behavioural vector spaces*

Trey's opening point was that most teams lean on ranking maths and hope it compensates for never really understanding the query. The talk was then a toolkit for fixing that upstream.

The headline idea is wormhole vectors. Standard hybrid search runs a lexical query and a dense query in parallel, then fuses the lists. A wormhole instead *travels* between spaces: run the lexical query, pool the dense vectors of the documents it matched (synthesising a new, pool-averaged vector), and search the dense space from that point. It also works in reverse, from dense results back into sparse terms via a [semantic knowledge graph](https://arxiv.org/abs/1609.00464). Around it he covered [late interaction](https://arxiv.org/abs/2004.12832) models and collaborative filtering for learning latent features from user behaviour.

Trey reports benchmarks beating conventional hybrid search (Wormhole precision@10 of 0.862 vs RRF precision@10 of 0.795). If that holds up broadly, it changes what a hybrid search API should expose: not just two queries and a fusion weight, but result-set pooling as a first-class operation. Another thing to add to your benchmarks and evaluations!

As both BM25 and dense search have got cheaper and faster, this concept is realistic in production with low double digit millisecond latency budgets. I would like to test it vs RRF with oversampling and semantic re-ranking.

### Sparse encoders: bridging products and knowledge graphs

[Event page](https://haystackconf.com/session/sparse-encoders-bridging-products-and-knowledge-graphs/) · [Slides (PDF)](https://pretalx.com/media/haystackeu26/submissions/MN9CDU/resources/Haystack_EU_2_cZZdGOm.pdf) · [Recording](https://www.youtube.com/watch?v=fwW0ujG_c74)

*[François Gaillard](https://www.linkedin.com/in/fran%C3%A7ois-gaillard-b2874988/), ADEO (Leroy Merlin, Bricoman)*

![sparse search for 'drill' showing weighted knowledge-graph concepts and top three matching drill products](https://simonhearne.com/images/haystack-europe-2026/sparse_encoders.png)
*A readable sparse vector: knowledge-graph concepts and weights for the query 'drill'*

ADEO runs home-improvement stores in Europe such as [Leroy Merlin](https://www.leroymerlin.fr/). ADEO's products have complex hierarchy and concepts (building materials vs. tools etc), so they need every product and query mapped onto its knowledge graph of concepts. Millions of product-concept pairs were poorly aligned, and a meaningful share of queries (18M / 21%) did not map correctly.

Dense embeddings didn't help: off-the-shelf models don't encode the same concepts, and the results couldn't be explained when a business rule misfired. This lack of 'explainability' is an issue I have encountered many times before, especially in e-commerce and legal document search scenarios.

François presented their solution to this challenge: a custom [SPLADE](https://arxiv.org/abs/2107.05720)-style sparse encoder which injects knowledge-graph concepts straight into the model vocabulary as tokens. Every sparse vector is then expressed in the catalogue's own vocabulary. A generative [T5 expansion](https://github.com/castorini/docTTTTTquery) approach lost on inference time (600 ms against 60 ms), but latency was reduced by adding [FLOPS-loss](https://arxiv.org/abs/2004.05665) and tuning the weights to push sparsity past 99.9%, without losing semantic depth.

This is the most practical case for learned sparse retrieval I've seen. You get semantic matching with concept understanding, a vector you can actually read and debug (it's just a list of terms), and it can be supported by any search engine that can index sparse vectors (such as [Milvus](https://milvus.io/docs/sparse_vector.md)).

### SPGE Enhanced Search: relevancy engine for Agentic RAG

[Event page](https://haystackconf.com/session/spge-enhanced-search-relevancy-engine-for-agentic-rag/) · [Slides (PDF)](https://pretalx.com/media/haystackeu26/submissions/7FLXGJ/resources/SPGE_Enhanced_4ejVhQL.pdf) · [Recording](https://www.youtube.com/watch?v=ydCBqb3VEFY)

*[Tomasz Florin](https://www.linkedin.com/in/tomasz-florin-93187553/) & [Artur Barczyński](https://www.linkedin.com/in/arturbarczynski/), S&P Global*

![diagram of one retrieval layer serving web apps, APIs, copilots and AI agents across relevance, experience, discoverability and AI](https://simonhearne.com/images/haystack-europe-2026/spge.png)
*One retrieval layer serving web apps, APIs, copilots and agents (S&P Global)*

S&P Global runs one hybrid relevancy platform for two very different consumers: the search box in its core enterprise product, and the retrieval layer behind agentic RAG workloads. Lexical and semantic retrieval are fused with RRF, and business-driven relevance rules shape the final ranking.

Tomasz and Artur did a good job of calling out the trade-offs in deploying a single retrieval stack. Complex business data needs precision and explainability as much as recall, and agents and analysts want different things from the same index. It's a clear picture of where large, regulated enterprises are heading: one retrieval platform, several consumers, with relevance rules the business can still control.

## New embedding models

### Hyperbolic embedding models for visual search

[Event page](https://haystackconf.com/session/hyperbolic-embedding-models-for-visual-search/) · [Recording](https://www.youtube.com/watch?v=qD-vNtcJpTc) · [Paper](https://arxiv.org/abs/2608.29313)

*[Matin Mahmood](https://www.linkedin.com/in/matin-mahmood/), [hyper³labs](https://hyper3labs.com/)*

![architecture diagram of an image-text model mapping captions and image regions into a hyperbolic Lorentz space](https://simonhearne.com/images/haystack-europe-2026/hyperbolic.png)
*Hyper3-CLIP embeds image regions and caption phrases in hyperbolic space*

Most image search stores hierarchy outside the vectors, in tags or a knowledge graph. hyper³labs trained an open image-text model ([Hyper3-CLIP](https://huggingface.co/hyper3labs/hyper3-clip-v1)) in [hyperbolic space](https://arxiv.org/abs/1705.08039) instead, where distance from the origin encodes specificity: "animal" sits near the centre, "border collie" near the edge. The hierarchy lives in the geometry, so a single nearest-neighbour query can respect it.

[The weights](https://huggingface.co/hyper3labs/hyper3-clip-v1) are open, and the results show good gains over the industry-standard [CLIP](https://arxiv.org/abs/2103.00020) multi-modal embedding model. The overall concept is intriguing and I can see many applications for it, from the obvious e-commerce category search to more interesting use cases like security incidents or fraud cases - where the specificity dimension could encode the maturity or depth of the attack / case.

I was so impressed with the concept that I have built it out on [Milvus](https://milvus.io/) - watch this space for the writeup!

### Introducing per-document embedding fine-tuning

[Event page](https://haystackconf.com/session/introducing-per-document-embedding-fine-tuning/) · [Recording](https://www.youtube.com/watch?v=4ERYs2F9B4A)

*[Philippe Bouzaglou](https://www.linkedin.com/in/philippe-bouzaglou-77227811b), Vectra*

![product search demo with sliders adjusting blue and red stone features on an earring embedding](https://simonhearne.com/images/haystack-europe-2026/composable_embeddings.png)
*Adjusting individual features of a product embedding at query time (Vectra)*

Embeddings are usually frozen at index time. They hold whatever the model could infer from the input, and changing them means fine-tuning the model and re-embedding the whole collection (which isn't cheap!). Vectra's e-commerce model is novel in that it has *deterministic dimensions* for attributes such as colour or material, so you can fine-tune an individual document's embedding at query time, without having to tune the model and re-generate the embeddings.

The demo that stuck with me: a product embedded from its photo, then enriched with facts the photo can't show, like organic cotton, a hippie look and festival wear. The same mechanism lets you tune embeddings in real time to a user's preferences.

If composable embeddings catch on, vectors stop being a static snapshot of a document and become something you update. That changes how we should think about index updates, and it's a very different workload from periodic bulk re-embedding.

## Evaluating search systems

### Beyond LLM Judges: Deep Evaluation for Conversational Search

[Event page](https://haystackconf.com/session/beyond-llm-judges-deep-evaluation-for-conversational-search/) · [Slides (PDF)](https://pretalx.com/media/haystackeu26/submissions/JDZWMV/resources/Beyond_LLM_Ju_eXkLm4A.pdf) · [Recording](https://www.youtube.com/watch?v=1E14GMCjFTQ)

*[Ruchi Juneja](https://www.linkedin.com/in/ruchi-juneja/), [MediaMarktSaturn](https://www.mediamarktsaturn.com/)*

![slide defining the Conversational Search Reliability Score with Excellent, Good, Fair and Poor bands](https://simonhearne.com/images/haystack-europe-2026/conversational_search.png)
*MediaMarktSaturn's Conversational Search Reliability Score*

MediaMarktSaturn's conversational shopping assistant had just gone into an A/B test, and Ruchi walked through why scoring only the final answer with an LLM judge wasn't enough. Their framework evaluates each layer separately: did the agent extract the right search terms, did its facets match the catalogue, how often did it hit zero results.

The standout finding was that implicit filters hurt. For example, if a customer mentioned "256GB", the agent turned it into a hard filter and narrowed the results, even though that shopper might have happily bought a 128GB or 512GB model. It's a small case of a big problem for agentic retrieval: an agent that follows the query too literally over-constrains it.

Evaluation is still largely human-graded on recall, accuracy and reliability, with business metrics not yet wired in. Her closing line is worth stealing: a good evaluation doesn't just tell you how good the system is, it tells you what to fix next.

Ruchi also described how they distilled their evaluation into a "Conversational Search Reliability Score" graded from Poor (< 50) to Excellent (≥ 90) so they could track progress and report to leadership. I love a simple metric like this, but I'm also curious whether the metric correlates with business metrics (conversion rate, abandonment etc.), which wasn't shared in the talk.

### Teaching your shopping assistant to learn from its mistakes

[Event page](https://haystackconf.com/session/teaching-your-shopping-assistant-to-learn-from-its-mistakes/) · [Slides (PDF)](https://pretalx.com/media/haystackeu26/submissions/R8ETJG/resources/HaystackEU26__JUFr2In.pdf) · [Recording](https://www.youtube.com/watch?v=iRytltsoC4E)

*[Jens Kürsten](https://www.linkedin.com/in/jenskuersten/), [OTTO](https://www.otto.de/)*

![slide showing session traces, shop interactions and conversations merged into a conversation timeline, beside the OTTO assistant app and the collect, judge, cluster, review, act loop](https://simonhearne.com/images/haystack-europe-2026/otto_assistant.png)
*Turning session traces into signals: the first step of OTTO's weekly loop*

This was a great companion to [Ruchi's talk](#beyond-llm-judges-deep-evaluation-for-conversational-search). When OTTO launched its product discovery assistant, failures were silent: shoppers just gave up. Jens described the weekly loop they now run over tens of thousands of production sessions:

1. **Collect** sessions and merge in the funnel events the customer actually saw.
2. **Judge** each one with an LLM anchored to those events.
3. **Cluster** failures into buckets with a single root cause each.
4. **Review** clusters against real dialogues to weed out LLM hallucinations.
5. **Act** on retrieval, prompt engineering or guardrails, then confirm in the next run.

The failures OTTO surfaced were interesting (and actionable): over-constrained queries returning zero results, gaps in catalogue attributes as well as purple team activity trying to break the guardrails. The key takeaway was to use LLMs to generate hypotheses and initial classification to improve the initial results, but use real customer interactions as well as human review to fine-tune in production.

## What I'm taking away

The centre of gravity in search has moved up the stack. Getting vectors into an index is assumed; the hard problems are understanding the query, serving agents and humans from one retrieval layer, and proving that the answer was right. The financial cost of all of this wasn't discussed at all outside my talk - but nothing is free. I feel like token costs and compute might become a limiting factor as these solutions scale.

I also have a lot of experiments and benchmarks to run!

