How to use from
SGLang
Install from pip and serve model
# Install SGLang from pip:
pip install sglang
# Start the SGLang server:
python3 -m sglang.launch_server \
    --model-path "cosmo3769/starcoderbase-1b-GPTQ" \
    --host 0.0.0.0 \
    --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "cosmo3769/starcoderbase-1b-GPTQ",
		"prompt": "Once upon a time,",
		"max_tokens": 512,
		"temperature": 0.5
	}'
Use Docker images
docker run --gpus all \
    --shm-size 32g \
    -p 30000:30000 \
    -v ~/.cache/huggingface:/root/.cache/huggingface \
    --env "HF_TOKEN=<secret>" \
    --ipc=host \
    lmsysorg/sglang:latest \
    python3 -m sglang.launch_server \
        --model-path "cosmo3769/starcoderbase-1b-GPTQ" \
        --host 0.0.0.0 \
        --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "cosmo3769/starcoderbase-1b-GPTQ",
		"prompt": "Once upon a time,",
		"max_tokens": 512,
		"temperature": 0.5
	}'
Quick Links

YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

starcoderbase-1b-GPTQ

Quantized starcoderbase-1b model to GPTQ format (4-bit precision) using Auto-GPTQ.

Quantization script

Benchmark

Benchmarking script

Baseline starcoderbase-1b model (non-quantized)

Tasks Version Filter n-shot Metric Value Stderr
codexglue_code2text N/A none None smoothed_bleu_4 0.8767 ± 0.0592
- code2text_go 1 none None smoothed_bleu_4 1.0054 ± 0.0983
- code2text_java 1 none None smoothed_bleu_4 1.2158 ± 0.1657
- code2text_javascript 1 none None smoothed_bleu_4 0.8560 ± 0.0429
- code2text_php 1 none None smoothed_bleu_4 0.9879 ± 0.0887
- code2text_python 1 none None smoothed_bleu_4 1.1950 ± 0.2819
- code2text_ruby 3 none None smoothed_bleu_4 0.0000 ± 0.0000
Groups Version Filter n-shot Metric Value Stderr
codexglue_code2text N/A none None smoothed_bleu_4 0.8767 ± 0.0592
Tasks Version Filter n-shot Metric Value Stderr
bigbench_code_line_description_generate_until 1 none None exact_match 0 ± 0
Tasks Version Filter n-shot Metric Value Stderr
bigbench_code_line_description_multiple_choice 0 none None acc 0.15 ± 0.0465

Quantized starcoderbase-1b model to GPTQ format

Tasks Version Filter n-shot Metric Value Stderr
codexglue_code2text N/A none None smoothed_bleu_4 0.7959 ± 0.2180
- code2text_go 1 none None smoothed_bleu_4 0.9280 ± 0.0291
- code2text_java 1 none None smoothed_bleu_4 1.2112 ± 0.1703
- code2text_javascript 1 none None smoothed_bleu_4 0.8848 ± 0.0391
- code2text_php 1 none None smoothed_bleu_4 0.6055 ± 0.6055
- code2text_python 1 none None smoothed_bleu_4 1.1460 ± 1.1460
- code2text_ruby 3 none None smoothed_bleu_4 0.0000 ± 0.0000
Groups Version Filter n-shot Metric Value Stderr
codexglue_code2text N/A none None smoothed_bleu_4 0.7959 ± 0.218
Tasks Version Filter n-shot Metric Value Stderr
bigbench_code_line_description_generate_until 1 none None exact_match 0 ± 0
Tasks Version Filter n-shot Metric Value Stderr
bigbench_code_line_description_multiple_choice 0 none None acc 0.1333 ± 0.0443
Downloads last month
43
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support