This project outlines a research plan to evaluate and compare two large language models for their effectiveness in analyzing student competence in Python programming.
The primary goal is to assess specialized code models on their ability to analyze student code submissions. The evaluation focuses on key criteria such as the ability to follow instructions, accurately diagnose misconceptions, provide pedagogical feedback without revealing solutions, and generate meaningful, non-revealing prompts.
This evaluation focuses on two specialized, 70-billion-parameter models from the CodeLlama family:
- CodeLlama-70B-Instruct: An instruction-tuned model designed for prompt-based analysis and assistant-like workflows. It is optimized to provide helpful and safe responses.(https://huggingface.co/codellama/CodeLlama-70b-Instruct-hf)
- CodeLlama-70B-Python: A model specifically tailored for understanding and handling Python-related tasks. It was further fine-tuned on 100B tokens of Python code.
- (https://huggingface.co/codellama/CodeLlama-70b-Python-hf)
The validation process is structured around four key areas:
- Method: Student code submissions will be assessed using unit testing to identify bugs and edge cases.
- Metric: The model's accuracy in pinpointing these errors will be measured.
- Method: Expert-defined rubrics will be used to evaluate the quality of the model's feedback. Meaningful prompt generation will be tested via expert blinded A/B evaluations against human-written prompts.
- Metrics:
- Conceptual coverage and the ability to address misconceptions.
- Avoidance of leaking complete solutions.
- Diagnostic sharpness and specificity of prompts to the submitted code.
- Method: The evaluation will adhere to responsible use guidelines to ensure a safe experience for users. Safety checks will follow guidance from the model card and practices recommended by Meta.
- Method: The performance and efficiency of the models will be checked under real-world conditions to assess serving costs and latency.
- CodeLlama-70B-Instruct was chosen as the primary engine because it is instruction-tuned for conversational and directive-following tasks, making it ideal for an interactive chatbot that can analyze code and generate prompts.
- CodeLlama-70B-Python was selected to complement the Instruct model with its deep, specialized understanding of Python code. This specialization makes it a strong backbone for deep code analysis or further task-specific fine-tuning.
- Accuracy vs. Cost: Larger, specialized models like the 70B variants provide higher accuracy but also have increased latency and serving costs. Smaller models are better suited for low-latency applications.
- Interpretability: Instruction-tuned models like
CodeLlama-70B-Instructoffer better out-of-the-box interpretability and the ability to generate safe explanations, whereas base models might require additional fine-tuning. - Context Window: Both evaluated models are fine-tuned for a context of up to 16,000 tokens and do not support 100k token contexts. This means that assessments of large code repositories would require chunking or retrieval augmentation strategies.
- Compute Requirements: The 70B scale of these models necessitates significant computational resources.