LLM · Text Generation
Meta Llama 3.1 8B OpenHermes 2.5
Llama 3.1 8B fine-tuned on OpenHermes 2.5 for instruction following and general language tasks.
About this model
This model is a fine-tuned version of Meta's Llama 3.1 8B, trained on the teknium/OpenHermes-2.5 dataset using BF16 mixed precision on a single NVIDIA A100 80GB GPU. The training ran for 13,368 steps with AdamW optimizer (learning rate ~2.5e-6), gradient accumulation of 8, and gradient checkpointing enabled. It uses the standard Llama 3.1 tokenizer (128,256 vocab) with 4,096 hidden size and 14,336 intermediate size. Licensed under Apache 2.0, it is intended for instruction following, question answering, and general language tasks. The model was trained using the Axolotl framework with Hugging Face Transformers.
Project signals
- 18 Hugging Face downloads
- 6 Hugging Face likes
Topics
transformers · safetensors · llama · text-generation · instruct · finetune · chatml · gpt4 · synthetic data · distillation