AI & ML interests
None defined yet.
Recent Activity
View all activity
Reality123bΒ
posted an update 14 days ago
Post
4562
We trained an open-source Mythos like cybersecurity LLM for the Build Small Hackathon meet OpenMythos
Trained in two stages: SFT on ~1.84K filtered ArXiv cs.CR papers + real CVE data, then RLVR using paired with past vulnerabilities GitHub repos with a verifier model checking outputs against ground truth.
Trained on: H100s from Modal
The RLVR stage made the biggest difference responses got more precise and less prone to confusing similar vulnerability classes.
Everything is open:
π€ Demo β build-small-hackathon/OpenMythos
π§ Model β build-small-hackathon/OpenMythos
π¦ CVE Dataset β build-small-hackathon/CVE_Vulnerailities_Detailed
π ArXiv Dataset β himanshu17HF/ArvixImport-Filtered-Final
Try it out and let us know where it breaks π
Trained in two stages: SFT on ~1.84K filtered ArXiv cs.CR papers + real CVE data, then RLVR using paired with past vulnerabilities GitHub repos with a verifier model checking outputs against ground truth.
Trained on: H100s from Modal
The RLVR stage made the biggest difference responses got more precise and less prone to confusing similar vulnerability classes.
Everything is open:
π€ Demo β build-small-hackathon/OpenMythos
π§ Model β build-small-hackathon/OpenMythos
π¦ CVE Dataset β build-small-hackathon/CVE_Vulnerailities_Detailed
π ArXiv Dataset β himanshu17HF/ArvixImport-Filtered-Final
Try it out and let us know where it breaks π
dippatel1994Β
posted an update about 2 months ago
Post
1101
To make revising LLM architectures and training methods faster, I created a deck of 180 visual flashcards. It started as a personal hobby, but slowly became cheat code for reviewing LLM concepts before technical interviews. People love it!
Swipe through these samples, and if you want to grab the full set or follow the project, the repo is here: https://github.com/llmsresearch/llm-flashcards.
Swipe through these samples, and if you want to grab the full set or follow the project, the repo is here: https://github.com/llmsresearch/llm-flashcards.
Pending Access Request to Join HF Blog Explorers
6
#17 opened 3 months ago
by
AINovice2005
alielfilali01Β
posted an update about 2 months ago
Post
647
Plans in HTML > Plans in Markdown
unmodeled-tylerΒ
posted an update 2 months ago
Post
203
Centauri ADE: https://github.com/unmodeled-tyler/centauri
This is Centauri - It's a lightweight Agent Development Environment for Linux with git source control and codebase understanding.
I've tried out several different IDE/ADEs and nothing has really worked as well as I'd like for what I need. For the options that I did like, BYOK was second-class, or in the case of a full IDE, it sometimes felt like I was trying to get an airliner into low-earth-orbit.
Enter: Centauri - a simple local-first CLI harness wrapper with a focused git source-control workspace.
It gives you a clean side-by-side flow:
- run your preferred CLI coding agent in an embedded terminal
- watch repository changes appear in the Changes panel
- generate an industry-standard commit message
- commit and push without leaving the app
It's designed to complement tools like Claude Code, Codex, Pi, OpenCode, etc. It does not replace Git, your terminal, or your existing credentials - it wraps the tools you already use into a tighter cohesive workspace optimized for engineering with an agent.
I built Centauri for myself after trying way too many different options. It solves several pain points for me and I'm sharing it in case it solves some for you too!
This is Centauri - It's a lightweight Agent Development Environment for Linux with git source control and codebase understanding.
I've tried out several different IDE/ADEs and nothing has really worked as well as I'd like for what I need. For the options that I did like, BYOK was second-class, or in the case of a full IDE, it sometimes felt like I was trying to get an airliner into low-earth-orbit.
Enter: Centauri - a simple local-first CLI harness wrapper with a focused git source-control workspace.
It gives you a clean side-by-side flow:
- run your preferred CLI coding agent in an embedded terminal
- watch repository changes appear in the Changes panel
- generate an industry-standard commit message
- commit and push without leaving the app
It's designed to complement tools like Claude Code, Codex, Pi, OpenCode, etc. It does not replace Git, your terminal, or your existing credentials - it wraps the tools you already use into a tighter cohesive workspace optimized for engineering with an agent.
I built Centauri for myself after trying way too many different options. It solves several pain points for me and I'm sharing it in case it solves some for you too!
rajkumarrawalΒ
submitted a
paper to Daily Papers 2 months ago
Post
3090
ππ»ββοΈ Hey there folks ,
Turns out : if we predict π earth we can save a lot of time looking for interesting things and less time looking at things that we expect to see.
Sentinel-2 imagery π°οΈbasically takes a long time to download towards earth. so our "near real time" systems are quite far from that in practical terms.
meanwhile , if we "predict" what we will see , based on what we do see , we can send down much less data in a timely way , and prioritize π‘earth-bound response .
I'm talking about illegal fishing , logging , mining or building in nature reserves , the more of that we predict early the more we're able to stop it on time.
At least that's the concept !
check out the blog : https://huggingface.co/blog/Tonic/save-patagonia-by-predicting-earth
- Collection: https://huggingface.co/collections/NuTonic/earth-observation-with-temporal-and-general-understanding
- Code: https://github.com/Josephrp/Nutonic
- Dataset: NuTonic/sat-vl-sft-training-ready-v1
- Model: NuTonic/lspace
- Training: NuTonic/lspace-trackio
- Evals: NuTonic/Patagonia_Eval
Turns out : if we predict π earth we can save a lot of time looking for interesting things and less time looking at things that we expect to see.
Sentinel-2 imagery π°οΈbasically takes a long time to download towards earth. so our "near real time" systems are quite far from that in practical terms.
meanwhile , if we "predict" what we will see , based on what we do see , we can send down much less data in a timely way , and prioritize π‘earth-bound response .
I'm talking about illegal fishing , logging , mining or building in nature reserves , the more of that we predict early the more we're able to stop it on time.
At least that's the concept !
check out the blog : https://huggingface.co/blog/Tonic/save-patagonia-by-predicting-earth
- Collection: https://huggingface.co/collections/NuTonic/earth-observation-with-temporal-and-general-understanding
- Code: https://github.com/Josephrp/Nutonic
- Dataset: NuTonic/sat-vl-sft-training-ready-v1
- Model: NuTonic/lspace
- Training: NuTonic/lspace-trackio
- Evals: NuTonic/Patagonia_Eval
unmodeled-tylerΒ
posted an update 2 months ago
Post
2983
The UFO/UAP Dataset is complete!
unmodeled-tyler/DoW-UFO-UAP-1
The most recent release from the Department of War is there up in full and ready for analysis!
The dataset ships with an Hermes Agent Skill so you can quickly and easily start parsing through the data immediately.
Go chase some anomalies! π
unmodeled-tyler/DoW-UFO-UAP-1
The most recent release from the Department of War is there up in full and ready for analysis!
The dataset ships with an Hermes Agent Skill so you can quickly and easily start parsing through the data immediately.
Go chase some anomalies! π
merveΒ
updated a
Space 2 months ago
rajkumarrawalΒ
posted an update 2 months ago
Post
2144
LLMs arenβt just answering questions anymore, theyβre learning to evolve. Self evolving AI is the true endgame.
AI has shifted from short tasks to long missions. The breakthrough isnβt just automation, itβs machines learning human methods and applying them at machine speed. From cybersecurity to finance, from OPCs to NPCs, the wave is irreversible.
Read the full article: Self Evolving is the Endgame or final destiny
https://huggingface.co/blog/rajkumarrawal/self-evolving-is-the-endgame-or-final-destiny
Whatβs your definition of true AGI? Comment below.
AI has shifted from short tasks to long missions. The breakthrough isnβt just automation, itβs machines learning human methods and applying them at machine speed. From cybersecurity to finance, from OPCs to NPCs, the wave is irreversible.
Read the full article: Self Evolving is the Endgame or final destiny
https://huggingface.co/blog/rajkumarrawal/self-evolving-is-the-endgame-or-final-destiny
Whatβs your definition of true AGI? Comment below.
[Support] Community Articles
ππ€ 2
106
#5 opened over 2 years ago
by
victor
[Support] Community Articles
π€π 2
106
#5 opened over 2 years ago
by
victor
unmodeled-tylerΒ
posted an update 3 months ago
Post
3017
Just started a fun project!
unmodeled-tyler/DoW-UFO-UAP-1
I'm getting the recently released DoW UFO/UAP documents (https://war.gov/ufo) cleaned and converted into a dataset here on Hugging Face!
There 161 different files in the gov release (pdfs, images, videos, audio, etc) and my current plan is to do it all in 1 dataset with 4 different shards - that way you can just call whichever tables you want/need when you import the dataset.
This is an ongoing project (I'm doing it on the side + my regular projects) so it's a bit of a growing entity. I'll also continuously refine the data over time to make sure it's as clean as possible.
Check it out! Who knows what you'll find in there?
unmodeled-tyler/DoW-UFO-UAP-1
I'm getting the recently released DoW UFO/UAP documents (https://war.gov/ufo) cleaned and converted into a dataset here on Hugging Face!
There 161 different files in the gov release (pdfs, images, videos, audio, etc) and my current plan is to do it all in 1 dataset with 4 different shards - that way you can just call whichever tables you want/need when you import the dataset.
This is an ongoing project (I'm doing it on the side + my regular projects) so it's a bit of a growing entity. I'll also continuously refine the data over time to make sure it's as clean as possible.
Check it out! Who knows what you'll find in there?
unmodeled-tylerΒ
posted an update 3 months ago
Post
4130
Hey Hugging Face!
Repo: https://github.com/unmodeled-tyler/vessel-browser
I wanted to share a cool feature from my open source AI native web browser, Vessel: Persistent highlights!
You can highlight anything on the page and the context is provided to the agent. It's kind of a fun way to learn about new stuff, synthesize info, or just deepen your comprehension/understanding.
Since highlights are persistent, you can close the page, come back later - and your highlights will be exactly where you left them. I've found this particularly useful when reviewing technical blogs, model cards, etc.
Check it out!
Repo: https://github.com/unmodeled-tyler/vessel-browser
I wanted to share a cool feature from my open source AI native web browser, Vessel: Persistent highlights!
You can highlight anything on the page and the context is provided to the agent. It's kind of a fun way to learn about new stuff, synthesize info, or just deepen your comprehension/understanding.
Since highlights are persistent, you can close the page, come back later - and your highlights will be exactly where you left them. I've found this particularly useful when reviewing technical blogs, model cards, etc.
Check it out!
CRAFTFrameworkΒ
posted an update 3 months ago
Post
117
# A measurable QA layer for LLM working sessions
Hallucination is treated as an inherent LLM failure mode, but most production workflows respond to it with "be careful, double-check things." That doesn't scale past a handful of sessions, and it doesn't catch the failure mode that does the most damage: confident reconstruction in late-session context.
CRAFT for Cowork takes a structural approach. The QA framework runs verification at four levels β individual claims, recipe execution, file integrity, and cross-session consistency β and treats trust as a measurable property rather than a vibe.
**The four-gate verification sub-routine** (RCP-CWK-024) runs before any recipe reports a result:
1. *File-pointability* β claim traceable to a specific file
2. *Read-vs-reconstructed* β was data actually read this session
3. *Lessons-Learned conflict* β contradicts documented prior truth
4. *Untested assumption* β verified vs. assumed
**Confidence scoring** grades every factual claim 0-100 against a source hierarchy: evidence read from files (80-100), tool-output observation (50-79), design intent (30-49), pure reasoning (0-29). A 10-point penalty applies past 70% token usage to correct for late-session reliability decay.
**Cross-session consistency** is enforced by a longitudinal audit recipe (RCP-CWK-036) run every 5-10 sessions. It has caught ~40% drift in tracking-file state tables β drift that would otherwise propagate as silent ground truth.
**Concrete result:** A factual claim validation pass caught nine pre-publication content files referencing the framework with an incorrect license descriptor. Single pass, all nine corrected.
This is week 5 of an 8-week capability spotlight. CRAFT for Cowork is a free public beta.
Repository: https://github.com/CRAFTFramework/craft-framework
License: https://craftframework.ai/craft-license/ (Spec under BSL 1.1, converts to Apache 2.0 on Jan 1, 2029; content proprietary)
Hallucination is treated as an inherent LLM failure mode, but most production workflows respond to it with "be careful, double-check things." That doesn't scale past a handful of sessions, and it doesn't catch the failure mode that does the most damage: confident reconstruction in late-session context.
CRAFT for Cowork takes a structural approach. The QA framework runs verification at four levels β individual claims, recipe execution, file integrity, and cross-session consistency β and treats trust as a measurable property rather than a vibe.
**The four-gate verification sub-routine** (RCP-CWK-024) runs before any recipe reports a result:
1. *File-pointability* β claim traceable to a specific file
2. *Read-vs-reconstructed* β was data actually read this session
3. *Lessons-Learned conflict* β contradicts documented prior truth
4. *Untested assumption* β verified vs. assumed
**Confidence scoring** grades every factual claim 0-100 against a source hierarchy: evidence read from files (80-100), tool-output observation (50-79), design intent (30-49), pure reasoning (0-29). A 10-point penalty applies past 70% token usage to correct for late-session reliability decay.
**Cross-session consistency** is enforced by a longitudinal audit recipe (RCP-CWK-036) run every 5-10 sessions. It has caught ~40% drift in tracking-file state tables β drift that would otherwise propagate as silent ground truth.
**Concrete result:** A factual claim validation pass caught nine pre-publication content files referencing the framework with an incorrect license descriptor. Single pass, all nine corrected.
This is week 5 of an 8-week capability spotlight. CRAFT for Cowork is a free public beta.
Repository: https://github.com/CRAFTFramework/craft-framework
License: https://craftframework.ai/craft-license/ (Spec under BSL 1.1, converts to Apache 2.0 on Jan 1, 2029; content proprietary)
rajkumarrawalΒ
posted an update 3 months ago
Post
231
I submitted a "Learning to Act and Cooperate for Distributed Black-Box Consensus Optimization" Paper by Zi-Bo Qin, Feng-Feng Wei, Tai-You Chen, Wei-Neng Chen to Daily Papers on huggingface.
A trajectory-driven framework uses large language models to guide agent behavior and cooperation patterns in distributed black-box consensus optimization, improving solution quality and efficiency.
Learning to Act and Cooperate for Distributed Black-Box Consensus Optimization (2605.00691)
A trajectory-driven framework uses large language models to guide agent behavior and cooperation patterns in distributed black-box consensus optimization, improving solution quality and efficiency.
Learning to Act and Cooperate for Distributed Black-Box Consensus Optimization (2605.00691)
rajkumarrawalΒ
submitted a
paper to Daily Papers 3 months ago
Future of Agentic Models
π₯ 2
12
#18 opened 3 months ago
by
MohamedRashad
Ujjwal-TyagiΒ
posted an update 3 months ago
Post
473
6 Open-Source Libraries to FineTune LLMs
1. Unsloth
GitHub: https://github.com/unslothai/unsloth
β Fastest way to fine-tune LLMs locally
β Optimized for low VRAM (even laptops)
β Plug-and-play with Hugging Face models
2. Axolotl
GitHub: https://github.com/OpenAccess-AI-Collective/axolotl
β Flexible LLM fine-tuning configs
β Supports LoRA, QLoRA, multi-GPU
β Great for custom training pipelines
3. TRL (Transformer Reinforcement Learning)
GitHub: https://github.com/huggingface/trl
β RLHF, DPO, PPO for LLM alignment
β Built on Hugging Face ecosystem
β Essential for post-training optimization
4. DeepSpeed
GitHub: https://github.com/microsoft/DeepSpeed
β Train massive models efficiently
β Memory + speed optimization
β Industry standard for scaling
5. LLaMA-Factory
GitHub: https://github.com/hiyouga/LLaMA-Factory
β All-in-one fine-tuning UI + CLI
β Supports multiple models (LLaMA, Qwen, etc.)
β Beginner-friendly + powerful
6. PEFT
GitHub: https://github.com/huggingface/peft
β Fine-tune with minimal compute
β LoRA, adapters, prefix tuning
β Best for cost-efficient training
1. Unsloth
GitHub: https://github.com/unslothai/unsloth
β Fastest way to fine-tune LLMs locally
β Optimized for low VRAM (even laptops)
β Plug-and-play with Hugging Face models
2. Axolotl
GitHub: https://github.com/OpenAccess-AI-Collective/axolotl
β Flexible LLM fine-tuning configs
β Supports LoRA, QLoRA, multi-GPU
β Great for custom training pipelines
3. TRL (Transformer Reinforcement Learning)
GitHub: https://github.com/huggingface/trl
β RLHF, DPO, PPO for LLM alignment
β Built on Hugging Face ecosystem
β Essential for post-training optimization
4. DeepSpeed
GitHub: https://github.com/microsoft/DeepSpeed
β Train massive models efficiently
β Memory + speed optimization
β Industry standard for scaling
5. LLaMA-Factory
GitHub: https://github.com/hiyouga/LLaMA-Factory
β All-in-one fine-tuning UI + CLI
β Supports multiple models (LLaMA, Qwen, etc.)
β Beginner-friendly + powerful
6. PEFT
GitHub: https://github.com/huggingface/peft
β Fine-tune with minimal compute
β LoRA, adapters, prefix tuning
β Best for cost-efficient training