CodeEmbed

Overview

CodeEmbed

The project explores neural code retrieval using transformer-based bi-encoder architectures, dense semantic retrieval, BM25 lexical retrieval, and hybrid retrieval strategies.

This repository contains trained model checkpoints from the CodeEmbed experiments, including baseline models, dual-encoder experiments, ablation studies, hybrid retrieval experiments, and Phase 7 evaluation models.

The CodeEmbed training experiments, retrieval system implementation, experiment configurations, evaluation workflow, and associated checkpoints were developed as part of this project.

Project Goals

CodeEmbed investigates effective retrieval of source-code functions from natural-language queries.

The project explores:

  • Transformer-based code representations
  • Bi-encoder architectures
  • Dense vector retrieval
  • BM25 lexical retrieval
  • Hybrid dense and lexical retrieval
  • Pooling strategies
  • Temperature experiments
  • Ablation studies
  • FAISS-based retrieval
  • Ranking-based retrieval evaluation

Dataset

The experiments use an AST-cleaned corpus derived from CodeSearchNet-based data.

The underlying dataset remains subject to its original licensing and attribution requirements.

Retrieval Approach

CodeEmbed supports hybrid retrieval by combining dense semantic retrieval, BM25 lexical retrieval, and convex interpolation of retrieval scores.

One demonstrated configuration uses:

  • Dense weight (alpha): 0.70
  • BM25 contribution: 0.30
  • Corpus size: 19,632

Checkpoints

The repository contains checkpoints from multiple CodeEmbed experiments.

Baseline and Dual Encoder Experiments

  • basic/
  • dual/
  • dual_bm25_hard/
  • dual_fixed/
  • dual_fixed_bm25/

Ablation Experiments

  • ablation/
  • ablation_basic_sweep/
  • ablation_dual_sweep/
  • ablation_dual_fixed_sweep/

These experiments investigate different pooling strategies and temperature values.

Shared Models

  • shared/
  • shared_large/

Phase 7 Models

The Phase 7 experiments include CodeBERT, Jina code model, MiniLM, and UniXcoder.

The checkpoints are located under phase7/.

Phase 7 Evaluation

A CodeEmbed Phase 7 CodeBERT experiment was evaluated on a test set containing 19,632 query-code pairs.

Metric Score
Test MRR 0.7658
Recall@1 0.6890
Recall@5 0.8640
Recall@10 0.9060
NDCG@10 0.7976

The reported validation MRR for the first epoch was 0.7356.

These results are specific to the project's evaluation setup and dataset.

Hybrid Retrieval Results

A demonstrated CodeEmbed hybrid retrieval configuration using convex interpolation achieved:

  • Test MRR: 0.6612
  • Improvement over the corresponding dense configuration: 40.7%
  • Search latency: approximately 216.81 ms
  • Corpus size: 19,632
  • Dense weight: 0.70
  • Observed leakage rate: 1.23%

These values are specific to the project's evaluation and demonstration configuration.

Training Checkpoints

The repository preserves both selected model checkpoints and intermediate training checkpoints.

Files named best_*.pt generally represent the selected checkpoint for an experiment.

Files named checkpoint_epoch_*.pt and checkpoint_step_*.pt represent intermediate training checkpoints.

Intermediate checkpoints are retained to preserve the experimental history.

Third-Party Models

Some experiments use or evaluate pretrained or externally developed models, including CodeBERT, UniXcoder, MiniLM, and Jina code models.

The inclusion of these checkpoints does not imply ownership of the underlying pretrained models, architectures, tokenizers, or original training data.

Third-party models and components remain subject to their respective licenses and terms.

Attribution and Provenance

CodeEmbed and the associated training experiments, retrieval implementation, evaluation workflow, and project-specific checkpoints were developed by Himasree Panku.

This Hugging Face repository preserves the trained checkpoint artifacts and provides a versioned record of the uploaded files.

Third-party models, datasets, libraries, and other external components remain subject to their original licenses and attribution requirements.

Reproducibility

For complete reproduction, users should obtain the corresponding CodeEmbed source code, configuration files, preprocessing pipeline, tokenizer and model configuration, and dependency environment.

Repository Structure

  • ablation/
  • ablation_basic_sweep/
  • ablation_dual_fixed_sweep/
  • ablation_dual_sweep/
  • basic/
  • dual/
  • dual_bm25_hard/
  • dual_fixed/
  • dual_fixed_bm25/
  • phase7/codebert/
  • phase7/jina_v2_code/
  • phase7/minilm_l6/
  • phase7/unixcoder/
  • shared/
  • shared_large/

License

The licensing of individual checkpoints may depend on the underlying pretrained models and third-party components used to produce them.

No license is granted here for third-party pretrained models, datasets, or other components beyond the rights provided by their respective licenses.

Before redistributing or licensing individual checkpoints, users should verify the applicable third-party licenses and permissions.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support