Back返回

CoDeR: Local Constraint-Compatible Retrieval Beyond Semantic Similarity

Published:

arXiv preprintarXiv 预印本 · 2026

CoDeR

CoDeR:超越语义相似度的局部约束相容检索CoDeR: Local Constraint-Compatible Retrieval Beyond Semantic Similarity

Xingkun Yin, Xuebin Tang, Hongyang Du

The Problem问题所在

Search engines and retrieval systems rank documents by how semantically similar they are to the query. For most queries that is a reasonable proxy for relevance. For some, it is actively wrong.搜索引擎与检索系统按文档与查询的语义相似度排序。对多数查询而言,这是相关性的合理代理指标;但对某些查询,它会直接给出错误答案。

Consider a query like "a quiet hotel for light sleep, a place away from noisy nightlife." The user is stating a constraint with a direction: quiet is what they want, noise is what they want to avoid. Now consider a review that reads "unlike other quiet hotels, this one is noisy at night." That review is topically extremely close to the query — same words, same subject — and it argues the exact opposite of what the user asked for. A similarity retriever will rank it highly.假设查询是「一家适合浅眠的安静酒店,远离吵闹的夜生活」。用户表达的是一条带方向的约束:安静是他要的,吵闹是他要避开的。再看一条评论:「和其他安静酒店不同,这家夜里很吵。」这条评论在主题上与查询极度接近 —— 同样的词、同样的对象 —— 却恰好论证了用户诉求的反面。而基于相似度的检索器会把它排得很靠前。

The paper names this failure constraint-violating evidence exposure. Similarity measures the topic; it has no notion of which direction the evidence points. So the retriever puts documents that violate the user's constraint into the top ranks, where the user actually looks.论文把这种失败命名为违约束证据暴露。相似度衡量的是「话题」,它不知道证据指向哪个方向。于是违反用户约束的文档被放进最靠前的位置 —— 也就是用户真正会看的地方。

Existing remedies sit at two extremes. Logic-based pipelines decompose the query into first-order logic, which needs translation or judging modules and adds latency and deployment complexity when implemented with strong external LLMs. Query-embedding optimisation is lightweight, but it only changes the query representation — the topical embedding space is left as it was, which limits how consistently it can separate satisfying from violating evidence.现有补救手段分处两个极端。基于逻辑的流程把查询拆解成一阶逻辑,需要翻译或判定模块;若用强外部大模型实现,会带来延迟与部署复杂度。查询向量优化则很轻量,但它改的只是查询表示 —— 主题嵌入空间原封不动,这限制了它稳定区分「满足约束」与「违反约束」证据的能力。

Key Idea核心思路

CoDeR separates two things that similarity collapses into one: topical relevance and constraint compatibility. A document can be highly relevant to the topic and still incompatible with the constraint.CoDeR 把相似度揉成一件事的两个概念拆开:主题相关性与约束兼容性。一篇文档可以主题高度相关,同时与约束并不兼容。

The insight核心洞见

Rather than rewriting the query to steer away from violating documents, put the pressure on the evidence itself: learn a representation in which constraint direction is separable from topical relatedness.与其改写查询去绕开违约束文档,不如把压力施加在证据本身:学出一个表示,让约束方向与主题关联性可分。

That means training a scorer on explicit satisfying-versus-violating contrasts, and applying it to the topically plausible candidates a normal retriever already returns. The intervention happens on the document side, locally, without any external LLM in the loop.具体做法是:用「满足 vs 违反」的显式对比训练一个打分器,再把它作用到普通检索器本来就返回的主题相关候选上。干预发生在文档侧、在本地,链路里没有任何外部大模型。

How It Works方法

CoDeR overview

Figure 1: CoDeR keeps a standard topical encoder for candidate coverage and adds a compatibility scorer trained on satisfying-versus-violating contrasts. The two signals are combined into a final ranking.图 1:CoDeR 保留标准的主题编码器以保证候选覆盖,另加一个用「满足 vs 违反」对比训练出的兼容性打分器。两路信号最终合成排序。

  1. Keep an ordinary topical retriever. A standard dense encoder still handles candidate coverage, so the system does not lose the recall that makes a retriever usable in the first place.保留一个普通的主题检索器。标准的稠密编码器仍负责候选召回,这样系统不会丢掉让检索器「能用」的那个前提 —— 召回率。
  2. Add a compatibility scorer. A second encoder — a bi-encoder initialised from bge-large-en-v1.5 — is trained to score constraint compatibility separately from topical similarity.增加一个兼容性打分器。第二个编码器 —— 由 bge-large-en-v1.5 初始化的双塔模型 —— 专门负责把约束兼容性与主题相似度分开打分。
  3. Train it on lexical polarity contrasts. Supervision comes in two stages: WordNet-derived word-level lexical polarity triplets, then sentence-level triplets drawn from 800 NevIR and 800 ExcluIR training queries. All queries, triplets and associated documents used to train the scorer are removed from the evaluation splits.用词汇极性对比来训练。监督信号分两阶段:先是由 WordNet 推导的词级词汇极性三元组,再是取自 800 条 NevIR 与 800 条 ExcluIR 训练查询的句级三元组。所有用于训练打分器的查询、三元组与相关文档都从评测集中剔除。
  4. Combine the two signals. The compatibility score can either rescore the topical candidates or retrieve an auxiliary compatibility-oriented candidate set, which is how the reported CoDeR-Seq and CoDeR-Union variants differ. Both produce a ranked list using only local encoder scoring — no external LLM calls at inference time.把两路信号合并。兼容性分数既可以用来对主题候选重新排序,也可以用来额外召回一组「面向兼容性」的候选 —— 这正是 CoDeR-Seq 与 CoDeR-Union 两个变体的区别。两者都只靠本地编码器打分生成排序,推理时不需要调用任何外部大模型。

Results实验结果

CoDeR is measured with two diagnostics that are specific to the failure it targets, rather than with general ranking quality alone:CoDeR 用的是针对该失败模式的两项专门指标,而不只是通用排序质量:

  • V@k — whether a violating document appears in the top k results (lower is better)V@k —— 前 k 条结果里是否出现了违反约束的文档(越低越好)
  • FVR — the rank of the first violating document (higher is better)FVR —— 第一条违约束文档所在的排名位置(越高越好)

It is evaluated on three controlled diagnostic sets, one per constraint type (antonymy, negation, exclusion), and on public negative-constraint retrieval benchmarks with stronger baselines including BM25, BGE, Contriever, HyDE, NS-IR and NevIR. All reported numbers are averaged over 10 independent runs and the ranking is stable across them.评测覆盖三组受控诊断集(分别对应反义、否定、排除三类约束),以及公开的负向约束检索基准,对比基线包括 BM25、BGE、Contriever、HyDE、NS-IR 与 NevIR。所有数字均为 10 次独立运行的平均值,排序在多次运行间保持稳定。

−20.59V@2 points on the antonymy diagnostic (73.53 → 52.94)反义诊断集上的 V@2 降幅(73.53 → 52.94)
−23.53V@2 points on the negation diagnostic (72.55 → 49.02)否定诊断集上的 V@2 降幅(72.55 → 49.02)
−5.77V@2 points on the exclusion diagnostic (53.85 → 48.08)排除诊断集上的 V@2 降幅(53.85 → 48.08)
0external LLM calls at inference time推理时需要的外部大模型调用次数

Table 1: Constraint-violating evidence in the top ranks. V@2 counts whether a violating document reaches the top two; the comparison is against the strongest non-CoDeR baseline on each diagnostic set.表 1:前排结果中的违约束证据。V@2 统计违约束文档是否进入前二;对照对象是各组诊断集上最强的非 CoDeR 基线。

DiagnosticStrongest non-CoDeR V@2CoDeR V@2Change
Antonymy73.5352.94−20.59
Negation72.5549.02−23.53
Exclusion53.8548.08−5.77

What the numbers say这些数字说明了什么

  1. The metric that matters is exposure, not average rank quality. A user reads the first few results. V@k asks whether a document that contradicts the stated constraint is sitting there, which is precisely the failure being studied — and it is invisible to an aggregate similarity score.真正要紧的指标是「暴露」,而不是平均排序质量。用户只看前几条。V@k 问的是:那条与明确约束相矛盾的文档,是不是就摆在那里 —— 这正是被研究的失败本身,而它在任何聚合相似度分数里都看不出来。
  2. Antonymy and negation are the hard cases. These are where a document keeps the query's vocabulary while flipping its meaning, so the topical signal actively misleads. The V@2 reductions there (20.6 and 23.5 points) are roughly four times the one on exclusion, where the violating documents are less lexically entangled with the query.反义与否定是最难的两类。这两种情况下,文档保留了查询的用词却把含义翻转,于是主题信号变成了误导。它们的 V@2 降幅(20.6 与 23.5)大约是排除类的四倍 —— 排除类里违约束文档与查询的词汇纠缠没那么深。
  3. General retrieval quality is preserved. The method is also evaluated on BEIR-style benchmarks with nDCG@10 and MAP@10, reported as a preservation check rather than a claim of general-purpose superiority: separating constraint direction should not cost ordinary topical coverage, and it does not.通用检索质量没有被牺牲。论文还在 BEIR 风格基准上报告了 nDCG@10 与 MAP@10 —— 这是「保真性检查」,不是宣称通用检索更强:把约束方向分出来,不应该以牺牲普通主题覆盖为代价,而事实也确实如此。
  4. The gains come from the representation, not from a cleverer prompt. Query-rewriting baselines make the query more informative but keep the inherited topical space. CoDeR instead trains on satisfying–violating contrasts, so the compatibility signal is available for any candidate set at no inference-time API cost.增益来自表示,而不是更聪明的提示词。查询改写类基线让查询携带更多信息,但继承的主题空间没变。CoDeR 则直接在「满足–违反」对比上训练,因此对任意候选集都能给出兼容性信号,且推理时零 API 成本。

Citation引用

If you find this work useful, please consider citing it:如果这项工作对您有帮助,欢迎引用:

@article{yin2026coder,
  title   = {CoDeR: Local Constraint-Compatible Retrieval Beyond Semantic Similarity},
  author  = {Yin, Xingkun and Tang, Xuebin and Du, Hongyang},
  journal = {arXiv preprint arXiv:2606.13204},
  year    = {2026}
}

The full manuscript is available as arXiv:2606.13204 (PDF).完整手稿见 arXiv:2606.13204(PDF)。