AI & ML interests

None defined yet.

Recent Activity

Yuki131  updated a collection 9 days ago
Lychee-JevEmbed
Yuki131  updated a model 9 days ago
HIT-TMG/JevEmbed-Qwen3-Embedding-4B
Yuki131  published a model 9 days ago
HIT-TMG/JevEmbed-Qwen3-Embedding-4B
View all activity

Yuki131 
posted an update 14 days ago
view post
Post
2821
Meet JevEmbed: an open-source framework for embedding-based decisions

Turn embeddings into decisions. Choose, score, and judge with your choice of embedding model.

We’ve open-sourced JevEmbed, a Python framework for three structured decision tasks:

🎯 Choice: select from a set of candidates
📊 Score: rate against ordered criteria
✅ Noul: judge whether a statement or question holds

🔧 JevEmbed currently includes configurations for KaLM, Qwen3, and E5 embedding models. You can use it through a Python API, CLI, or optional HTTP server. It also supports local LoRA fine-tuning, so you can adapt an embedding model to your own decision tasks and load the resulting adapter for local inference.

Fine-tuning results

📈 We trained KaLM-Embedding-V2.5 and Qwen3-Embedding-0.6B on the 79,116-example training split of Open-Jev’s release-v2-redistributable subset. We then evaluated them on 3,495 hard-label questions from the same subset’s held-out validation split.

ZefanCai/Open-Jev

KaLM-Embedding-V2.5: 30.24% base accuracy → 76.68% after LoRA fine-tuning
Qwen3-Embedding-0.6B: 30.73% base accuracy → 84.06% after LoRA fine-tuning

KaLM-Embedding/KaLM-embedding-multilingual-mini-instruct-v2.5
Qwen/Qwen3-Embedding-0.6B

These results are specific to that validation split. Performance on other tasks and datasets may differ.

JevEmbed also supports Choice tasks with more than 255 candidates, making it useful for classification and routing problems with large candidate sets.

Explore the framework, open an issue, or tell us what decision task you would try it on:
🔗 https://github.com/HITsz-TMG/JevEmbed

#Embeddings #LoRA #SentenceTransformers #OpenSource #JevEmbed
  • 3 replies
·
Yuki131 
posted an update 16 days ago
view post
Post
3657
Meet KaLM-Jev — your local, Jev-style judgment engine, available in Nano, Small, and Large.

Building an agent or automation workflow? Sometimes all you need is a choice, a score, or a signal that a condition holds.

Built on KaLM-Reranker-R2, KaLM-Jev turns these decisions into structured outputs through three primitives:

🔀 Choice — select among candidates, with a probability distribution.
📊 Score — return a continuous score over your defined levels.
🔍 Noul — evaluate conditions independently, so multiple conditions can hold at once.

Think support-ticket routing, bug severity scoring, human-escalation detection, or candidate tool selection for agents.

🖥️ Run locally with downloaded weights
📦 Choose from Nano / Small / Large
🔌 Integrate through HTTP or Python
⚡ Reuse cached candidate/rule representations to reduce repeated encoding
🧪 Explore included examples, bilingual semantic smoke tests, and recorded GPU validation results

No answer-text generation: output_tokens = 0. Inference still runs to compute the judgments.

KaLM-Jev is an independent implementation based on KaLM-Reranker, not an official TypeSafe project or a guarantee of full Jev compatibility. Scores are uncalibrated; validate thresholds on your own tasks.

Code & quickstart:
https://github.com/KaLM-Embedding/KaLM-Jev
Yuki131/KaLM-Jev
We’d love to hear what you’d build with it. Try it out, share feedback, or open an issue! 🤗

#Jev #Reranker #Agents #LocalAI #OpenSource
  • 1 reply
·
xyidealist 
in HIT-TMG/Lychee-FD about 1 month ago
HITSZ-TMG 
in HIT-TMG/Lychee-FD about 2 months ago

Move model files to repository root

#8 opened about 2 months ago by
xyidealist
xyidealist 
in HIT-TMG/Lychee-FD about 2 months ago

Move model files to repository root

#8 opened about 2 months ago by
xyidealist
Yuki131 
posted an update about 2 months ago
view post
Post
2264
Test-Time Scaling for Rerankers?

Can rerankers scale at test time—not by generating longer reasoning traces, but by selectively using richer document representations?


KaLM-Reranker-V1 supports Matryoshka compression from 1× to 32×, which suggests a progressive multi-fidelity pipeline:

- Embedding retrieval → Top-100
- KaLM-Reranker @ 32× compression → Top-20
- The same reranker @ 2× compression → final ranking

The intuition is simple: cheaply screen many candidates, then allocate higher-fidelity cross-attention only to the most promising ones.

For 100@32× → 20@2×, the passage-token interaction budget is roughly 31.8% of directly running 100@2×, before fixed model overheads. The key question is whether it can retain nearly the same ranking quality.

We’re considering evaluating nDCG–latency Pareto curves.

Would you consider this a useful form of test-time scaling for retrieval?

KaLM-Embedding/KaLM-Reranker-V1-Nano

KaLM-Embedding/KaLM-Reranker-V1-Small

KaLM-Embedding/KaLM-Reranker-V1-Large

KaLM-Reranker-V1: Fast but Not Late Interaction for Compressed Document Reranking (2606.22807)

https://huggingface.co/collections/KaLM-Embedding/lychee-kalm-reranker

KaLM-Embedding
  • 2 replies
·