Skip to content

About

A plan to evaluate CodeLlama-70B-Instruct and CodeLlama-70B-Python for analyzing student Python competence. The models are assessed on correctness, pedagogy, safety, and cost.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

4 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 

Repository files navigation

Evaluating Open Source Models for Student Competence Analysis

This project outlines a research plan to evaluate and compare two large language models for their effectiveness in analyzing student competence in Python programming.


Objective

The primary goal is to assess specialized code models on their ability to analyze student code submissions. The evaluation focuses on key criteria such as the ability to follow instructions, accurately diagnose misconceptions, provide pedagogical feedback without revealing solutions, and generate meaningful, non-revealing prompts.


Models Under Evaluation

This evaluation focuses on two specialized, 70-billion-parameter models from the CodeLlama family:

  1. CodeLlama-70B-Instruct: An instruction-tuned model designed for prompt-based analysis and assistant-like workflows. It is optimized to provide helpful and safe responses.(https://huggingface.co/codellama/CodeLlama-70b-Instruct-hf)
  2. CodeLlama-70B-Python: A model specifically tailored for understanding and handling Python-related tasks. It was further fine-tuned on 100B tokens of Python code.
  3. (https://huggingface.co/codellama/CodeLlama-70b-Python-hf)

Evaluation Framework

The validation process is structured around four key areas:

1. Correctness

  • Method: Student code submissions will be assessed using unit testing to identify bugs and edge cases.
  • Metric: The model's accuracy in pinpointing these errors will be measured.

2. Pedagogy

  • Method: Expert-defined rubrics will be used to evaluate the quality of the model's feedback. Meaningful prompt generation will be tested via expert blinded A/B evaluations against human-written prompts.
  • Metrics:
    • Conceptual coverage and the ability to address misconceptions.
    • Avoidance of leaking complete solutions.
    • Diagnostic sharpness and specificity of prompts to the submitted code.

3. Safety

  • Method: The evaluation will adhere to responsible use guidelines to ensure a safe experience for users. Safety checks will follow guidance from the model card and practices recommended by Meta.

4. Cost

  • Method: The performance and efficiency of the models will be checked under real-world conditions to assess serving costs and latency.

Justification for Model Selection

  • CodeLlama-70B-Instruct was chosen as the primary engine because it is instruction-tuned for conversational and directive-following tasks, making it ideal for an interactive chatbot that can analyze code and generate prompts.
  • CodeLlama-70B-Python was selected to complement the Instruct model with its deep, specialized understanding of Python code. This specialization makes it a strong backbone for deep code analysis or further task-specific fine-tuning.

Known Trade-offs and Limitations

  • Accuracy vs. Cost: Larger, specialized models like the 70B variants provide higher accuracy but also have increased latency and serving costs. Smaller models are better suited for low-latency applications.
  • Interpretability: Instruction-tuned models like CodeLlama-70B-Instruct offer better out-of-the-box interpretability and the ability to generate safe explanations, whereas base models might require additional fine-tuning.
  • Context Window: Both evaluated models are fine-tuned for a context of up to 16,000 tokens and do not support 100k token contexts. This means that assessments of large code repositories would require chunking or retrieval augmentation strategies.
  • Compute Requirements: The 70B scale of these models necessitates significant computational resources.

About

A plan to evaluate CodeLlama-70B-Instruct and CodeLlama-70B-Python for analyzing student Python competence. The models are assessed on correctness, pedagogy, safety, and cost.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors