DeBERTa-v3-large-direct-metaphor
DeBERTa-v3-large fine-tuned for MIPVU direct metaphor detection (token-level, 3-class).
Identifies metaphor flag words (mFlag) and source-domain content words (mrw_lit) in English text.
Model Description
This model detects direct metaphors as defined by the MIPVU procedure (Steen et al., 2010). A direct metaphor occurs when words are used literally but the overall expression introduces a conceptually incongruous source domain through an explicit comparison signal (mFlag), creating a cross-domain mapping.
Examples:
| Sentence | mFlag | Source-domain words |
|---|---|---|
| "like a buzzard in its eyrie" | like | buzzard, eyrie |
| "treated the workers as pawns" | as | pawns |
| "the sword-shaped cloud" | -shaped | sword |
Label Scheme (token-level, 3 classes)
| ID | Label | Description |
|---|---|---|
| 0 | O |
No annotation β non-metaphor token |
| 1 | mFlag |
Metaphor flag: the comparison signal word (e.g. like, as, regarded as, -shaped) |
| 2 | mrw_lit |
MRW-direct: source-domain content word used literally within the direct metaphor span |
Intended Use
- Primary: Token-level direct metaphor detection in English text following MIPVU annotation standards
- Downstream: Sentence-level screening for direct metaphors (sent F1 = 83.4% on VUAMC validation)
- Pipeline role: Can be combined with an indirect metaphor model (e.g.
deberta-v3-large-metaphor) for full MIPVU-style metaphor annotation
Not intended for:
- Languages other than English
- Indirect metaphor detection (see companion model)
- Implicit metaphor detection
Training Data
Training corpus: BE06-SpaCy (British English, 500 texts)
- 1,114 positive sentences with LLM-generated mFlag / mrw_lit annotations
- 2,228 hard-negative sentences (contain mFlag-lexicon words but no direct metaphor;
neg_ratio=2) - Total: 3,342 sentences, ~300K tokens
- Token distribution: O=98.5%, mFlag=0.64%, mrw_lit=0.86%
Validation corpus: VUAMC-SpaCy (gold standard)
- 110 positive sentences with VUAMC mFlag / mrw_lit annotations
- 200 hard-negative sentences (stratified by trigger word: as 51%, like 26%, similar 8%, other 15%)
Training annotations note: The BE06 training labels were generated by a large language model applying the MIPVU direct metaphor procedure. The validation set uses human gold-standard VUAMC annotations, providing an out-of-distribution quality signal.
Training Procedure
| Hyperparameter | Value |
|---|---|
| Base model | microsoft/deberta-v3-large |
| Epochs | 4 (best at epoch 2) |
| Learning rate | 2e-5 |
| LR scheduler | Linear with warmup |
| Warmup steps | 84 (10% of total) |
| Effective batch size | 16 |
| Max sequence length | 256 |
| Weight decay | 0.01 |
| Class weights | O=1.0, mFlag=5.0, mrw_lit=5.0 (sqrt-inv-freq, capped at 5.0) |
| Optimizer | AdamW |
| Precision | float32 |
| Best checkpoint metric | Combined token F1 (mFlag + mrw_lit as positive) |
Class weighting was applied to compensate for the severe positive-class imbalance (~1.5% combined positive token rate). The cap of 5.0 prevents gradient instability.
Evaluation Results
Evaluated on VUAMC validation set (110 positive + 200 hard-negative sentences):
Token-level
| Class | Precision | Recall | F1 |
|---|---|---|---|
| mFlag | 69.18% | 76.92% | 72.85% |
| mrw_lit | 74.84% | 78.33% | 76.55% |
| Combined | 73.36% | 78.33% | 75.76% |
Sentence-level
| Metric | Value |
|---|---|
| Precision | 75.18% |
| Recall | 93.64% |
| F1 | 83.40% |
Comparison with LLM Baseline
| DeBERTa (this model) | DeepSeek-V4-Flash (zero-shot MIPVU) | |
|---|---|---|
| sent F1 | 83.40% | 87.50% |
| mFlag F1 | 72.85% | ~95.9% (acc) |
| lit F1 | 76.55% | 86.0% |
The model achieves competitive sentence-level detection (β4.1pp vs LLM) with a large speed advantage. The token-level gap reflects the LLM's access to explicit MIPVU symbolic rules.
Usage
from transformers import AutoTokenizer, AutoModelForTokenClassification
import torch
model_name = "path/to/deberta-v3-large-direct-metaphor"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForTokenClassification.from_pretrained(model_name)
model.eval()
# Pre-tokenised word list
words = ["He", "fought", "like", "a", "lion", "in", "battle"]
inputs = tokenizer(
words,
is_split_into_words=True,
return_tensors="pt",
truncation=True,
max_length=256,
)
with torch.no_grad():
logits = model(**inputs).logits # (1, seq_len, 3)
# Align predictions back to word level (first subword only)
word_ids = inputs.word_ids(batch_index=0)
id2label = model.config.id2label
seen, word_preds = set(), []
for tok_i, wid in enumerate(word_ids):
if wid is None or wid in seen:
continue
seen.add(wid)
label_id = logits[0, tok_i].argmax().item()
word_preds.append((words[wid], id2label[str(label_id)]))
for word, label in word_preds:
print(f"{word:12s} β {label}")
# He β O
# fought β O
# like β mFlag
# a β O
# lion β mrw_lit
# in β O
# battle β O
Limitations
- Validation set size: Only 310 validation sentences β per-epoch estimates have moderate variance
- Training data noise: BE06 annotations are LLM-generated, not human gold standard
- Token span precision: Source-domain span boundary detection is imprecise compared to rule-based systems
- mFlag categories: Coverage of less common mFlag types (Category D mental-framing verbs, Category E metalinguistic adverbs) may be limited by training data frequency
- Domain: Trained on British English (BE06 / VUAMC); may generalise less well to other varieties
Citation
Model author: Tommy Leo β 1683619168tl@gmail.com
If you use this model, please cite the MIPVU procedure:
@book{steen2010method,
title = {A Method for Linguistic Metaphor Identification:
From MIP to MIPVU},
author = {Steen, Gerard J. and Dorst, Aletta G. and Herrmann,
J. Berenike and Kaal, Anna and Krennmayr, Tina and
Pasma, Trijntje},
year = {2010},
publisher = {John Benjamins},
address = {Amsterdam}
}
License
Apache License 2.0 β see LICENSE for details.
- Downloads last month
- 7
Model tree for tommyleo2077/deberta-v3-large-direct-metaphor
Base model
microsoft/deberta-v3-large