Instructions to use dball/zephyr-7b-dpo-qlora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use dball/zephyr-7b-dpo-qlora with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("mistralai/Mistral-7B-v0.1") model = PeftModel.from_pretrained(base_model, "dball/zephyr-7b-dpo-qlora") - Notebooks
- Google Colab
- Kaggle
Is the high drop in GSM8K usual?
#1
by dball - opened
Overall comparison shows that most metrics improve through DPO:
dball/zephyr-7b-dpo-qlora: (+)Average: 61.27; -GSM8K 33.97; (+)Win 78.61; +TQA: 44.03; -+MMLU 62.28;+ HSwag 84.92; +ARC 63.82
dball/zephyr-7b-sft-qlora (-)Average: 59.8; -GSM8k 34.12; (-)Win 78.22; (+)TQA 42.32; -MMLU 61.9; (-)HSwag 82.49; (-)ARC 59.73
mistralai/Mistral-7B-v0.1: Average: 60.97; GSM8K 37.83; Win 78.37; TQA: 42.15; MMLU 64.16; HSwag 83.31; ARC 59.98