DeBERTa-v3-large-direct-metaphor

DeBERTa-v3-large fine-tuned for MIPVU direct metaphor detection (token-level, 3-class).
Identifies metaphor flag words (mFlag) and source-domain content words (mrw_lit) in English text.


Model Description

This model detects direct metaphors as defined by the MIPVU procedure (Steen et al., 2010). A direct metaphor occurs when words are used literally but the overall expression introduces a conceptually incongruous source domain through an explicit comparison signal (mFlag), creating a cross-domain mapping.

Examples:

Sentence mFlag Source-domain words
"like a buzzard in its eyrie" like buzzard, eyrie
"treated the workers as pawns" as pawns
"the sword-shaped cloud" -shaped sword

Label Scheme (token-level, 3 classes)

ID Label Description
0 O No annotation β€” non-metaphor token
1 mFlag Metaphor flag: the comparison signal word (e.g. like, as, regarded as, -shaped)
2 mrw_lit MRW-direct: source-domain content word used literally within the direct metaphor span

Intended Use

  • Primary: Token-level direct metaphor detection in English text following MIPVU annotation standards
  • Downstream: Sentence-level screening for direct metaphors (sent F1 = 83.4% on VUAMC validation)
  • Pipeline role: Can be combined with an indirect metaphor model (e.g. deberta-v3-large-metaphor) for full MIPVU-style metaphor annotation

Not intended for:

  • Languages other than English
  • Indirect metaphor detection (see companion model)
  • Implicit metaphor detection

Training Data

Training corpus: BE06-SpaCy (British English, 500 texts)

  • 1,114 positive sentences with LLM-generated mFlag / mrw_lit annotations
  • 2,228 hard-negative sentences (contain mFlag-lexicon words but no direct metaphor; neg_ratio=2)
  • Total: 3,342 sentences, ~300K tokens
  • Token distribution: O=98.5%, mFlag=0.64%, mrw_lit=0.86%

Validation corpus: VUAMC-SpaCy (gold standard)

  • 110 positive sentences with VUAMC mFlag / mrw_lit annotations
  • 200 hard-negative sentences (stratified by trigger word: as 51%, like 26%, similar 8%, other 15%)

Training annotations note: The BE06 training labels were generated by a large language model applying the MIPVU direct metaphor procedure. The validation set uses human gold-standard VUAMC annotations, providing an out-of-distribution quality signal.


Training Procedure

Hyperparameter Value
Base model microsoft/deberta-v3-large
Epochs 4 (best at epoch 2)
Learning rate 2e-5
LR scheduler Linear with warmup
Warmup steps 84 (10% of total)
Effective batch size 16
Max sequence length 256
Weight decay 0.01
Class weights O=1.0, mFlag=5.0, mrw_lit=5.0 (sqrt-inv-freq, capped at 5.0)
Optimizer AdamW
Precision float32
Best checkpoint metric Combined token F1 (mFlag + mrw_lit as positive)

Class weighting was applied to compensate for the severe positive-class imbalance (~1.5% combined positive token rate). The cap of 5.0 prevents gradient instability.


Evaluation Results

Evaluated on VUAMC validation set (110 positive + 200 hard-negative sentences):

Token-level

Class Precision Recall F1
mFlag 69.18% 76.92% 72.85%
mrw_lit 74.84% 78.33% 76.55%
Combined 73.36% 78.33% 75.76%

Sentence-level

Metric Value
Precision 75.18%
Recall 93.64%
F1 83.40%

Comparison with LLM Baseline

DeBERTa (this model) DeepSeek-V4-Flash (zero-shot MIPVU)
sent F1 83.40% 87.50%
mFlag F1 72.85% ~95.9% (acc)
lit F1 76.55% 86.0%

The model achieves competitive sentence-level detection (βˆ’4.1pp vs LLM) with a large speed advantage. The token-level gap reflects the LLM's access to explicit MIPVU symbolic rules.


Usage

from transformers import AutoTokenizer, AutoModelForTokenClassification
import torch

model_name = "path/to/deberta-v3-large-direct-metaphor"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForTokenClassification.from_pretrained(model_name)
model.eval()

# Pre-tokenised word list
words = ["He", "fought", "like", "a", "lion", "in", "battle"]

inputs = tokenizer(
    words,
    is_split_into_words=True,
    return_tensors="pt",
    truncation=True,
    max_length=256,
)

with torch.no_grad():
    logits = model(**inputs).logits  # (1, seq_len, 3)

# Align predictions back to word level (first subword only)
word_ids = inputs.word_ids(batch_index=0)
id2label = model.config.id2label
seen, word_preds = set(), []
for tok_i, wid in enumerate(word_ids):
    if wid is None or wid in seen:
        continue
    seen.add(wid)
    label_id = logits[0, tok_i].argmax().item()
    word_preds.append((words[wid], id2label[str(label_id)]))

for word, label in word_preds:
    print(f"{word:12s} β†’ {label}")
# He           β†’ O
# fought       β†’ O
# like         β†’ mFlag
# a            β†’ O
# lion         β†’ mrw_lit
# in           β†’ O
# battle       β†’ O

Limitations

  • Validation set size: Only 310 validation sentences β€” per-epoch estimates have moderate variance
  • Training data noise: BE06 annotations are LLM-generated, not human gold standard
  • Token span precision: Source-domain span boundary detection is imprecise compared to rule-based systems
  • mFlag categories: Coverage of less common mFlag types (Category D mental-framing verbs, Category E metalinguistic adverbs) may be limited by training data frequency
  • Domain: Trained on British English (BE06 / VUAMC); may generalise less well to other varieties

Citation

Model author: Tommy Leo β€” 1683619168tl@gmail.com

If you use this model, please cite the MIPVU procedure:

@book{steen2010method,
  title     = {A Method for Linguistic Metaphor Identification:
               From MIP to MIPVU},
  author    = {Steen, Gerard J. and Dorst, Aletta G. and Herrmann,
               J. Berenike and Kaal, Anna and Krennmayr, Tina and
               Pasma, Trijntje},
  year      = {2010},
  publisher = {John Benjamins},
  address   = {Amsterdam}
}

License

Apache License 2.0 β€” see LICENSE for details.

Downloads last month
7
Safetensors
Model size
0.4B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for tommyleo2077/deberta-v3-large-direct-metaphor

Finetuned
(315)
this model

Collection including tommyleo2077/deberta-v3-large-direct-metaphor