MoAI-Privacy-Filter-INT8

MoAI-Privacy-Filter-INT8 is an INT8 weight-only ONNX Runtime artifact for Korean and English entity detection. It identifies 29 entity types with 117 BIOES token classes.

This repository contains the quantized ONNX graph, external tensor data, tokenizer, label configuration, taxonomy, and Viterbi calibration. It does not contain PyTorch weights and cannot be loaded with AutoModelForTokenClassification.from_pretrained().

1. Model summary

  • Base model: openai/privacy-filter.
  • Parent checkpoint: BCCard/MoAI-Privacy-Filter v3 final-bf16.
  • Task: Korean and English token classification with character-offset entity reconstruction.
  • Taxonomy: 29 entity labels and 117 BIOES classes under 4N+1.
  • Input policy: Up to 1024 tokens per sequence.
  • Quantization: INT8 weight-only with FP32 activations and logits.
  • Runtime: ONNX Runtime CPU execution.

The validated serving chain is:

text
-> tokenizer
-> INT8 weight-only ONNX graph
-> FP32 logits
-> constrained BIOES Viterbi decoding
-> character-offset spans
-> whitespace boundary refinement

2. Labels

The model detects the following entity types:

Label Description
PERSON Person name
RRN Korean resident registration number
FRN Korean foreign resident registration number
SSN Social Security number
GENERIC_ID Identity or tax identifier without a more specific taxonomy class
CARD_NUMBER Payment card number
ACCOUNT_NUMBER Financial account number
SECRET Password, API key, token, or other authentication secret
USER_ID User or account login identifier
EMAIL Email address
PHONE Telephone number
PASSPORT Passport number
DRIVER_LICENSE Driver's license number
ADDRESS Postal or street address
ZIPCODE Postal code
DATE Date or time expression
CARD_EXPIRY Payment card expiration date
CVC Payment card verification code
IPIN Korean I-PIN identifier
TRANSACTION_APPROVAL_ID Transaction approval identifier
BUSINESS_ID Business registration identifier
VIRTUAL_CARD_NUMBER Virtual card number
CI Connecting Information identifier
IPADDRESS IP address
MACADDRESS MAC address
IMEI Mobile equipment identifier
PORT Network port number
ORGANIZATION Organization name
URL URL

Each entity has B-, I-, E-, and S- boundary classes. The remaining class is O.

3. Usage

Install the required packages:

pip install "onnxruntime>=1.28,<1.29" "huggingface-hub>=1.5" "transformers>=5.6" numpy

The following example runs the graph and returns FP32 logits:

import json
from pathlib import Path

import numpy as np
import onnxruntime as ort
from huggingface_hub import snapshot_download
from transformers import AutoTokenizer

model_id = "BCCard/MoAI-Privacy-Filter-INT8"
model_dir = Path(snapshot_download(repo_id=model_id))

tokenizer = AutoTokenizer.from_pretrained(model_dir)
session = ort.InferenceSession(
    str(model_dir / "model_quantized.onnx"),
    providers=["CPUExecutionProvider"],
)

text = "์—ฐ๋ฝ์ฒ˜๋Š” 010-1234-5678์ด๊ณ  ์ ‘์† ์ฃผ์†Œ๋Š” 192.0.2.15์ž…๋‹ˆ๋‹ค."
encoded = tokenizer(
    text,
    add_special_tokens=False,
    truncation=True,
    max_length=1024,
    return_tensors="np",
)
feeds = {
    model_input.name: np.asarray(encoded[model_input.name], dtype=np.int64)
    for model_input in session.get_inputs()
}
logits = session.run(["logits"], feeds)[0]

config = json.loads((model_dir / "config.json").read_text(encoding="utf-8"))
id2label = {
    int(class_id): label
    for class_id, label in config["id2label"].items()
}

print(logits.shape)
print(id2label[int(logits[0, 0].argmax())])

The graph emits logits rather than final entities. Apply the constrained BIOES Viterbi decoder using config.json and viterbi_calibration.json, then map token predictions to character offsets. Independent token argmax is not equivalent to the decoding chain used for validation.

4. Evaluation

The INT8 artifact and its FP32 ONNX reference were evaluated on the same 14,524-row Korean and English validation split. Metrics use strict character-span matching after constrained BIOES Viterbi decoding and whitespace boundary refinement.

Slice Model Precision Recall F1 F1 delta from FP32
Overall FP32 ONNX 0.9604 0.9597 0.9601 -
Overall INT8 0.9603 0.9596 0.9599 -0.0001
Korean FP32 ONNX 0.9586 0.9576 0.9581 -
Korean INT8 0.9584 0.9575 0.9580 -0.0001
English FP32 ONNX 0.9654 0.9654 0.9654 -
English INT8 0.9652 0.9654 0.9653 -0.0001

The INT8 predictions achieved 0.9984 strict micro F1 against the FP32 ONNX predictions, with 14,402 of 14,524 rows matching exactly. All configured quantization quality gates passed.

These end-to-end character-span scores are not directly comparable with training-time token metrics because they include character reconstruction, constrained decoding, and boundary refinement.

5. Artifact size

Artifact Graph files Relative size
BF16 parent checkpoint 2.799 GB Reference
INT8 weight-only ONNX 1.618 GB 42.2% smaller than the BF16 parent

The comparison uses the publicly released BF16 parent checkpoint as its reference. Actual memory use and latency depend on hardware, ONNX Runtime version, sequence length, and batch size.

6. Intended use and limitations

This model is intended for entity detection in privacy filtering, data review, and preprocessing workflows. A consuming application must define its own downstream redaction or retention policy for every detected label.

  • Evaluate the model on representative in-domain data before deployment.
  • Inputs longer than 1024 tokens require chunking and span reconciliation.
  • Context-dependent and ambiguous identifiers can still produce false positives or false negatives.
  • PORT and ZIPCODE, ORGANIZATION and PERSON, and structurally similar numeric identifiers require particular monitoring.
  • Do not treat model output as a substitute for organizational privacy controls or human review in high-risk workflows.

7. Training data and attribution

The parent model was trained with BCCard/privacy-filter-openpii-masking, which is derived in part from ai4privacy/pii-masking-openpii-1.5m and supplemented with Korean and English synthetic scenarios.

The base model is openai/privacy-filter. Review the licenses and usage terms of the model, dataset, and dependencies before redistribution or deployment.

8. License, Attribution, and Citation

@misc{bccard2026moaiprivacyfilter,
  title        = {MoAI Privacy Filter INT8: A Korean Finance-Domain PII Detection Model},
  author       = {BC Card AX Team},
  year         = {2026},
  howpublished = {https://huggingface.co/BCCard/MoAI-Privacy-Filter-INT8},
  note         = {INT8 weight-only ONNX artifact of a full fine-tune of openai/privacy-filter}
}

Related resources:

Downloads last month
65
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for BCCard/MoAI-Privacy-Filter-INT8

Quantized
(8)
this model

Dataset used to train BCCard/MoAI-Privacy-Filter-INT8

Collection including BCCard/MoAI-Privacy-Filter-INT8