LLM
Qwen2.5-0.5B Synthia II GGUF
GGUF quantizations of a Qwen2.5-0.5B model fine-tuned on Synthia-v1.5-II for instruction following and chat, with 32K context and ChatML template.
About this model
This model provides GGUF-format quantizations of Qwen2.5-0.5B-Synthia-II, a fine-tuned version of Qwen/Qwen2.5-0.5B trained on the Synthia-v1.5-II dataset for conversational AI and instruction following. The base model has 490M parameters (360M non-embedding) with 24 transformer layers and GQA attention (14 query heads, 2 key/value heads). It supports 32,768 token context length via RoPE embeddings, SwiGLU activations, and RMSNorm.
The fine-tuning used 3 epochs with a 1e-5 learning rate, batch size 40 (5×8 gradient accumulation), cosine LR schedule with 100 warmup steps, 4096 sequence length with sample packing, and BF16 mixed precision. Training completed in 672 steps. The model uses ChatML format with tool-calling support in its chat template.
Available quantizations: Q2_K, Q3_K_L, Q3_K_M, Q3_K_S, Q4_0, Q4_K_M, Q4_K_S, Q5_0, Q5_K_M, Q5_K_S, Q6_K, Q8_0, and f16. These run efficiently on CPU and Apple Silicon via llama.cpp and compatible tools. Licensed Apache-2.0.
Project signals
- 399 Hugging Face downloads
- 2 Hugging Face likes
Topics
transformers · gguf · instruct · finetune · chatml · gpt4 · synthetic data · distillation · dataset:migtissera/synthia-v1.5-ii · endpoints_compatible