BGE Small Structural Separator

This is the complete Transformers checkpoint for the H216 Shared Structural Separator Field Encoder. It is compatible with AutoModel and AutoTokenizer.

H216 starts from the immutable BAAI/bge-small-en-v1.5@5c38ec7c405ec4b44b94cc5a9bb96e735b38267a checkpoint and learns exactly one 384-value input-embedding row: token [unused2], vocabulary ID 3. All other model parameters are unchanged.

Document formatting inserts [unused2] before a nonempty title and before each punctuation-delimited sentence. The document is encoded once and its normalized CLS vector is stored. Queries use ordinary tokenizer formatting with no instruction. Retrieval is exact cosine over one query vector and one document vector; there is no fusion, routing, reranking, expansion, or protected candidate frontier.

from transformers import AutoModel, AutoTokenizer

model_id = "thu-nmrc/bge-small-structural-separator"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModel.from_pretrained(model_id)

The repository includes separator_row.pt, the deterministic training manifest, and export_manifest.json with SHA-256 hashes for the complete checkpoint. Training used 4,096 S2ORC citation-context/abstract pairs and no BEIR labels, validation selection, benchmark negative mining, or post-training tuning.

The bounded research claim is a learned document-side structural separator adaptation within the dense bi-encoder family. It is not claimed as a replacement for all dense retrievers. Full evaluation details, training code, overlap audits, and disclosed regressions are maintained at https://github.com/thu-nmrc/bge-small-structural-separator.

The frozen BEIR8 gate improves macro NDCG@10 by +0.012705 over the matched plain BGE control. A later preregistered six-task MTEB(eng, v2) Retrieval extension does not support broad transfer: macro NDCG@10 changes by -0.011038 and H216 wins two of six tasks. Across all ten official Retrieval tasks the macro delta is +0.002855, but claims remain bounded to the BEIR8 structure-native evidence.

Downloads last month
22
Safetensors
Model size
33.4M params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for thu-nmrc/bge-small-structural-separator

Finetuned
(372)
this model

Space using thu-nmrc/bge-small-structural-separator 1