AI & ML interests

None defined yet.

Recent Activity

Reality123bΒ 
posted an update 14 days ago
view post
Post
125
can you people please benchmark this thing i made? (i mean benchmark as in benchmarks) here is the link: xylaria2s.vercel.app/api/v1/chat/completions and please give benchmarks plus proofs underneath this post, please help me out
  • 2 replies
Β·
KingNishΒ 
posted an update about 1 month ago
view post
Post
4562
We trained an open-source Mythos like cybersecurity LLM for the Build Small Hackathon meet OpenMythos

Trained in two stages: SFT on ~1.84K filtered ArXiv cs.CR papers + real CVE data, then RLVR using paired with past vulnerabilities GitHub repos with a verifier model checking outputs against ground truth.

Trained on: H100s from Modal

The RLVR stage made the biggest difference responses got more precise and less prone to confusing similar vulnerability classes.

Everything is open:
πŸ€– Demo β†’ build-small-hackathon/OpenMythos
🧠 Model β†’ build-small-hackathon/OpenMythos
πŸ“¦ CVE Dataset β†’ build-small-hackathon/CVE_Vulnerailities_Detailed
πŸ“„ ArXiv Dataset β†’ himanshu17HF/ArvixImport-Filtered-Final

Try it out and let us know where it breaks πŸ™
  • 2 replies
Β·
dippatel1994Β 
posted an update about 2 months ago
view post
Post
1101
To make revising LLM architectures and training methods faster, I created a deck of 180 visual flashcards. It started as a personal hobby, but slowly became cheat code for reviewing LLM concepts before technical interviews. People love it!

Swipe through these samples, and if you want to grab the full set or follow the project, the repo is here: https://github.com/llmsresearch/llm-flashcards.
alielfilali01Β 
posted an update about 2 months ago
view post
Post
647
Plans in HTML > Plans in Markdown
unmodeled-tylerΒ 
posted an update 2 months ago
view post
Post
203
Centauri ADE: https://github.com/unmodeled-tyler/centauri

This is Centauri - It's a lightweight Agent Development Environment for Linux with git source control and codebase understanding.

I've tried out several different IDE/ADEs and nothing has really worked as well as I'd like for what I need. For the options that I did like, BYOK was second-class, or in the case of a full IDE, it sometimes felt like I was trying to get an airliner into low-earth-orbit.

Enter: Centauri - a simple local-first CLI harness wrapper with a focused git source-control workspace.

It gives you a clean side-by-side flow:
- run your preferred CLI coding agent in an embedded terminal
- watch repository changes appear in the Changes panel
- generate an industry-standard commit message
- commit and push without leaving the app

It's designed to complement tools like Claude Code, Codex, Pi, OpenCode, etc. It does not replace Git, your terminal, or your existing credentials - it wraps the tools you already use into a tighter cohesive workspace optimized for engineering with an agent.

I built Centauri for myself after trying way too many different options. It solves several pain points for me and I'm sharing it in case it solves some for you too!
  • 1 reply
Β·
TonicΒ 
posted an update 2 months ago
view post
Post
3090
πŸ™‹πŸ»β€β™‚οΈ Hey there folks ,

Turns out : if we predict 🌏 earth we can save a lot of time looking for interesting things and less time looking at things that we expect to see.

Sentinel-2 imagery πŸ›°οΈbasically takes a long time to download towards earth. so our "near real time" systems are quite far from that in practical terms.

meanwhile , if we "predict" what we will see , based on what we do see , we can send down much less data in a timely way , and prioritize πŸ“‘earth-bound response .

I'm talking about illegal fishing , logging , mining or building in nature reserves , the more of that we predict early the more we're able to stop it on time.

At least that's the concept !

check out the blog : https://huggingface.co/blog/Tonic/save-patagonia-by-predicting-earth


- Collection: https://huggingface.co/collections/NuTonic/earth-observation-with-temporal-and-general-understanding
- Code: https://github.com/Josephrp/Nutonic
- Dataset: NuTonic/sat-vl-sft-training-ready-v1
- Model: NuTonic/lspace
- Training: NuTonic/lspace-trackio
- Evals: NuTonic/Patagonia_Eval
  • 2 replies
Β·
unmodeled-tylerΒ 
posted an update 2 months ago
view post
Post
2983
The UFO/UAP Dataset is complete!

unmodeled-tyler/DoW-UFO-UAP-1

The most recent release from the Department of War is there up in full and ready for analysis!

The dataset ships with an Hermes Agent Skill so you can quickly and easily start parsing through the data immediately.

Go chase some anomalies! πŸš€

merveΒ 
updated a Space 2 months ago
rajkumarrawalΒ 
posted an update 2 months ago
view post
Post
2144
LLMs aren’t just answering questions anymore, they’re learning to evolve. Self evolving AI is the true endgame.

AI has shifted from short tasks to long missions. The breakthrough isn’t just automation, it’s machines learning human methods and applying them at machine speed. From cybersecurity to finance, from OPCs to NPCs, the wave is irreversible.

Read the full article: Self Evolving is the Endgame or final destiny

https://huggingface.co/blog/rajkumarrawal/self-evolving-is-the-endgame-or-final-destiny

What’s your definition of true AGI? Comment below.
  • 1 reply
Β·
unmodeled-tylerΒ 
posted an update 3 months ago
view post
Post
3017
Just started a fun project!

unmodeled-tyler/DoW-UFO-UAP-1

I'm getting the recently released DoW UFO/UAP documents (https://war.gov/ufo) cleaned and converted into a dataset here on Hugging Face!

There 161 different files in the gov release (pdfs, images, videos, audio, etc) and my current plan is to do it all in 1 dataset with 4 different shards - that way you can just call whichever tables you want/need when you import the dataset.

This is an ongoing project (I'm doing it on the side + my regular projects) so it's a bit of a growing entity. I'll also continuously refine the data over time to make sure it's as clean as possible.

Check it out! Who knows what you'll find in there?
  • 3 replies
Β·
unmodeled-tylerΒ 
posted an update 3 months ago
view post
Post
4130
Hey Hugging Face!

Repo: https://github.com/unmodeled-tyler/vessel-browser

I wanted to share a cool feature from my open source AI native web browser, Vessel: Persistent highlights!

You can highlight anything on the page and the context is provided to the agent. It's kind of a fun way to learn about new stuff, synthesize info, or just deepen your comprehension/understanding.

Since highlights are persistent, you can close the page, come back later - and your highlights will be exactly where you left them. I've found this particularly useful when reviewing technical blogs, model cards, etc.

Check it out!
  • 1 reply
Β·
CRAFTFrameworkΒ 
posted an update 3 months ago
view post
Post
117
# A measurable QA layer for LLM working sessions

Hallucination is treated as an inherent LLM failure mode, but most production workflows respond to it with "be careful, double-check things." That doesn't scale past a handful of sessions, and it doesn't catch the failure mode that does the most damage: confident reconstruction in late-session context.

CRAFT for Cowork takes a structural approach. The QA framework runs verification at four levels β€” individual claims, recipe execution, file integrity, and cross-session consistency β€” and treats trust as a measurable property rather than a vibe.

**The four-gate verification sub-routine** (RCP-CWK-024) runs before any recipe reports a result:

1. *File-pointability* β€” claim traceable to a specific file
2. *Read-vs-reconstructed* β€” was data actually read this session
3. *Lessons-Learned conflict* β€” contradicts documented prior truth
4. *Untested assumption* β€” verified vs. assumed

**Confidence scoring** grades every factual claim 0-100 against a source hierarchy: evidence read from files (80-100), tool-output observation (50-79), design intent (30-49), pure reasoning (0-29). A 10-point penalty applies past 70% token usage to correct for late-session reliability decay.

**Cross-session consistency** is enforced by a longitudinal audit recipe (RCP-CWK-036) run every 5-10 sessions. It has caught ~40% drift in tracking-file state tables β€” drift that would otherwise propagate as silent ground truth.

**Concrete result:** A factual claim validation pass caught nine pre-publication content files referencing the framework with an incorrect license descriptor. Single pass, all nine corrected.

This is week 5 of an 8-week capability spotlight. CRAFT for Cowork is a free public beta.

Repository: https://github.com/CRAFTFramework/craft-framework
License: https://craftframework.ai/craft-license/ (Spec under BSL 1.1, converts to Apache 2.0 on Jan 1, 2029; content proprietary)
  • 1 reply
Β·
rajkumarrawalΒ 
posted an update 3 months ago
view post
Post
231
I submitted a "Learning to Act and Cooperate for Distributed Black-Box Consensus Optimization" Paper by Zi-Bo Qin, Feng-Feng Wei, Tai-You Chen, Wei-Neng Chen to Daily Papers on huggingface.

A trajectory-driven framework uses large language models to guide agent behavior and cooperation patterns in distributed black-box consensus optimization, improving solution quality and efficiency.

Learning to Act and Cooperate for Distributed Black-Box Consensus Optimization (2605.00691)
Ujjwal-TyagiΒ 
posted an update 3 months ago
view post
Post
473
6 Open-Source Libraries to FineTune LLMs
1. Unsloth
GitHub: https://github.com/unslothai/unsloth
β†’ Fastest way to fine-tune LLMs locally
β†’ Optimized for low VRAM (even laptops)
β†’ Plug-and-play with Hugging Face models

2. Axolotl
GitHub: https://github.com/OpenAccess-AI-Collective/axolotl
β†’ Flexible LLM fine-tuning configs
β†’ Supports LoRA, QLoRA, multi-GPU
β†’ Great for custom training pipelines

3. TRL (Transformer Reinforcement Learning)
GitHub: https://github.com/huggingface/trl
β†’ RLHF, DPO, PPO for LLM alignment
β†’ Built on Hugging Face ecosystem
β†’ Essential for post-training optimization

4. DeepSpeed
GitHub: https://github.com/microsoft/DeepSpeed
β†’ Train massive models efficiently
β†’ Memory + speed optimization
β†’ Industry standard for scaling

5. LLaMA-Factory
GitHub: https://github.com/hiyouga/LLaMA-Factory
β†’ All-in-one fine-tuning UI + CLI
β†’ Supports multiple models (LLaMA, Qwen, etc.)
β†’ Beginner-friendly + powerful

6. PEFT
GitHub: https://github.com/huggingface/peft
β†’ Fine-tune with minimal compute
β†’ LoRA, adapters, prefix tuning
β†’ Best for cost-efficient training
  • 1 reply
Β·