N
LLM FINETUNING AND DATA INTERN
Norar
-
Internship
0-3 Years experience
Description
This is a remote position.
We are seeking a motivated and technically strong LLM Fine-Tuning and Data Engineering Intern to support our AI initiatives focused on large language model training and optimization.
The selected candidate will assist in preparing high-quality datasets, structuring instruction-tuning formats, and supporting fine-tuning workflows using modern open-source LLM frameworks.
This internship is ideal for candidates who want hands-on exposure to real-world LLM training pipelines, data curation processes, and model evaluation techniques.
Requirements
- Data Preparation and Cleaning
- Clean, normalize, and pre process raw text datasets
- Remove duplicate, irrelevant, or low-quality entries
- Structure datasets into JSON/JSONL formats suitable for training
- Prepare instruction–response formatted datasets
- Maintain proper train/validation splits
- Tokenization and Data Structuring
- Apply appropriate tokenizer for specific base models
- Analyze token length distributions
- Implement truncation and padding strategies
- Ensure dataset compatibility with sequence length constraints
- Support Fine-Tuning Workflows
- Assist in setting up LoRA / QLoRA / PEFT-based fine-tuning
- Configure training parameters under supervision
- Monitor training metrics such as loss curves
- Maintain organized experiment records
- Evaluation and Quality Assurance
- Conduct prompt-based evaluation of fine-tuned models
- Compare baseline vs fine-tuned model outputs
- Document findings and suggest improvements
- Documentation and Reporting
- Maintain clear documentation of datasets and experiments
- Prepare structured reports on data quality and training outcomes
- Ensure reproducibility of experiments
Benefits
Internship Details
- Duration: 3–6 months
- Mode: Remote / Hybrid (as applicable)
- Stipend: As per industry standards and candidate capability
- Opportunity for full-time role based on performance
About Norar
-
