Instructions to use Abhayn01/ARKA-SLM-V1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Abhayn01/ARKA-SLM-V1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Abhayn01/ARKA-SLM-V1") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Abhayn01/ARKA-SLM-V1") model = AutoModelForCausalLM.from_pretrained("Abhayn01/ARKA-SLM-V1", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Abhayn01/ARKA-SLM-V1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Abhayn01/ARKA-SLM-V1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Abhayn01/ARKA-SLM-V1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Abhayn01/ARKA-SLM-V1
- SGLang
How to use Abhayn01/ARKA-SLM-V1 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Abhayn01/ARKA-SLM-V1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Abhayn01/ARKA-SLM-V1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Abhayn01/ARKA-SLM-V1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Abhayn01/ARKA-SLM-V1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Abhayn01/ARKA-SLM-V1 with Docker Model Runner:
docker model run hf.co/Abhayn01/ARKA-SLM-V1
- ARKA SLM V1
- Model Information
- About ARKA
- Release Checkpoint
- Weight Integrity
- Fine-Tuning Recommendation
- Retrieval-Augmented Generation (RAG)
- Mobile and Edge Applications
- Prompt Format
- Python Usage
- Current Strengths
- Known Limitations
- Memory
- Recommended Production Architecture
- High-Stakes Use
- Creator
- Version
- License
- Model Information
ARKA SLM V1
ARKA SLM V1 is the first official release of the ARKA Small Language Model family.
ARKA stands for Academic Research Knowledge Assistant.
Creator and Developer: Abhay Kumar Rudrapaul
ARKA SLM V1 is a compact causal language model designed as a foundation for conversational AI, educational applications, RAG systems, mobile AI experiments and domain-specific assistants.
Fine-tuning is recommended before specialized or production deployment.
Retrieval-Augmented Generation (RAG) is strongly recommended for factual, private, domain-specific or frequently changing information.
Model Information
| Property | Value |
|---|---|
| Official Name | ARKA SLM V1 |
| Full Form | Academic Research Knowledge Assistant |
| Creator | Abhay Kumar Rudrapaul |
| Version | V1 |
| Release Status | Official Release |
| Model Type | Small Language Model |
| Architecture | GPT2LMHeadModel |
| Parameters | 129,944,832 |
| Hidden Size | 768 |
| Transformer Layers | 12 |
| Attention Heads | 12 |
| Head Dimension | 64 |
| Vocabulary Size | 50,257 |
| Maximum Context Positions | 8,192 |
| Framework | PyTorch / Hugging Face Transformers |
About ARKA
ARKA is a Small Language Model project created and developed by Abhay Kumar Rudrapaul.
The project explores compact language models combined with supervised fine-tuning, instruction tuning, retrieval, external tools and application-specific AI.
ARKA SLM V1 is intended to provide a compact foundation that developers can further adapt for their own use cases.
Release Checkpoint
- Source repository:
Abhayn01/arka-v5-8k-mixed-sft-54m-v1 - Source revision:
a4a4f1566c57f5bb2ce3417c184fc7ffc94a6ca5 - Source stage:
balanced_recovery_2p5m_v1 - Source status:
complete - Source step:
unknown - Recorded cumulative valid-token exposure: 76,539,609
- Recorded cumulative supervised-token exposure: unknown
The official release does not contain optimizer, gradient-scaler or intermediate training checkpoint state.
Weight Integrity
- NaN values: 0
- Infinite values: 0
- Fully-zero parameter tensors: 0
- Tied LM head / token embeddings: True
- Weight integrity: Verified
All model parameters were set to requires_grad=False during the official release-build verification.
Important: requires_grad=False is a runtime PyTorch property and is not permanently stored in SafeTensors. Users can intentionally fine-tune their own loaded copy.
Fine-Tuning Recommendation
ARKA SLM V1 should generally be treated as a foundation checkpoint.
Fine-tuning is recommended before using the model for a specialized domain or production application.
Examples include:
- educational assistants
- electrical engineering assistants
- campus assistants
- customer-support systems
- mobile personal assistants
- domain-specific question answering
- document assistants
- application intent routing
Retrieval-Augmented Generation (RAG)
RAG is strongly recommended when ARKA is expected to answer factual or knowledge-intensive questions.
Typical pipeline:
User Question
|
v
Intent Router
|
+---- Direct task ----> ARKA
|
+---- Knowledge task -> Retriever
|
v
Retrieved Evidence
|
v
ARKA SLM V1
|
v
Grounded Response
Possible retrieval sources include PDFs, local documents, vector databases, verified websites, institutional data, company documentation and search engines.
RAG can improve grounding, but it does not guarantee perfect factual accuracy.
Mobile and Edge Applications
ARKA SLM V1 contains approximately 129.9 million parameters.
It may be integrated into mobile or local applications using an appropriate runtime and, where necessary, quantization or model conversion.
Potential applications include:
- mobile AI assistants
- educational applications
- campus assistants
- voice assistants
- offline or partially offline assistants
- document Q&A systems
- domain-specific RAG applications
- embedded conversational features
Actual performance depends on model precision, RAM, CPU/GPU/NPU capability, runtime and context length.
Prompt Format
User: <user message>
Assistant:
Python Usage
from transformers import AutoTokenizer, AutoModelForCausalLM
MODEL_ID = "Abhayn01/ARKA-SLM-V1"
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)
model = AutoModelForCausalLM.from_pretrained(MODEL_ID)
prompt = "User: Explain renewable energy simply.\nAssistant:"
inputs = tokenizer(prompt, return_tensors="pt")
output = model.generate(
**inputs,
max_new_tokens=120,
do_sample=False,
eos_token_id=tokenizer.eos_token_id,
pad_token_id=tokenizer.eos_token_id,
)
new_tokens = output[0, inputs['input_ids'].shape[1]:]
print(tokenizer.decode(new_tokens, skip_special_tokens=True))
Current Strengths
Development testing has shown useful behavior in:
- ARKA identity handling
- short instruction following
- exact short responses
- supplied-context question answering
- explicit missing-context abstention
- some follow-up corrections
- short-context recall
- RAG-oriented workflows
These observations are developmental and are not standardized benchmark claims.
Known Limitations
ARKA SLM V1 is a compact model and has important limitations.
- factual recall can be unreliable
- unsupported questions may produce hallucinations
- arithmetic reasoning is weak
- programming ability is limited
- electrical knowledge is inconsistent
- multi-entity context binding can fail
- long-form text may become repetitive
- structured email/story/dialogue generation is inconsistent
- exact formatting or units may occasionally be lost
- reasoning capability is below significantly larger models
For factual tasks, developers should prefer RAG and validate important outputs.
Memory
ARKA SLM V1 should not be assumed to remember information from previous conversations unless the application explicitly provides that information in the active context.
Persistent memory should be implemented at the application layer.
Recommended Production Architecture
+--> Direct ARKA
|
User -> Intent Router+--> RAG -> ARKA
|
+--> Calculator / Tools
|
+--> Application Actions
High-Stakes Use
ARKA SLM V1 should not be used as the sole authority for medical, legal, financial, emergency, safety-critical or similarly high-stakes decisions.
Creator
Abhay Kumar Rudrapaul
Creator and developer of the ARKA model family.
Version
ARKA SLM V1 — Version 1.0
This checkpoint represents the first officially designated release of the ARKA Small Language Model family.
Future ARKA releases may improve factual knowledge, reasoning, long-form generation, RAG, tool use, coding, domain specialization and efficient mobile inference.
License
No specific open-source license is automatically declared by this release script. The repository owner should select an appropriate license separately before defining redistribution or commercial-use rights.
- Downloads last month
- 748