Hands on with Buckets

Community Article
Published September 1, 2026

A few years ago at a startup I built with some friends (TripLingo), one of my colleagues added our massive language translation datasets into our main Git repo. As you can guess, this wasn’t the best thing to do. Now all of a sudden all git operations were very slow - whether you git clone, fetch, or checkout. Of course, a solution was later invented called Git Large File Storage (LFS), which uses a pointer to the file and stores it in a separate server meant for large files. Git is inherently an immutable data storage mechanism, which means that files can’t be changed once committed (you can actually change a committed file, but it’s cumbersome) - these files have to be completely replaced.

This means Git LFS doesn’t work well for many of the large files we use in the AI world during development, whether it’s model checkpoints, conversation traces, inference outputs, etc, or any datasets that are still being added to, or worked on, and changing. And using S3 storage isn’t optimal either: it has enterprise level complexity and is slow. Hence, not a good option while you’re doing active development… developers want things to be FAST so we can get velocity and finish projects sooner.

So what’s a better option?

Enter a technology called Xet and Huggingface Buckets. Xet breaks files into 64KB chunks, so when you upload the same large file, the client calculates what’s changed, and only uploads the difference. Buckets allows you to easily and intelligently sync a file or a directory of files and it lets you save large amounts of bandwidth (and time and cost!) as well as storage. You can read more details about Buckets on the Huggingface blog about the cost savings, but in this tutorial we’ll get hands-on with Xet and Buckets so you can follow along and see it in action for yourself!

First Steps with Xet and Buckets

Let’s create a small project to demonstrate the Xet capabilities and how to use Huggingface Buckets. Our goal is to learn the basics of these two using a simple, easy to understand example. We’ll use this data in upcoming articles to do more interesting things, but for now, here’s what we’ll do in this simple demo project:

  1. Download a large dataset file which is in CSV format
  2. Log into Huggingface
  3. Create a Bucket
  4. Copy a file into the Bucket
  5. Append some data to the file
  6. Copy the file again and see how quickly it uploads compared to previously!
  7. Convert the file to a Parquet fileUse the sync command to sync a folder to the bucket with our new parquet file

Let’s get started!

First, let’s create our project directory:

mkdir simple-buckets-demo
cd simple-buckets-demo

Then, go to the link here and click the download link on the top right. Save the file in the project directory. The file is about 200M in size and is a standard CSV file. Next we need to install the Huggingface CLI and login. Have a look at the CLI guide if you don’t have these two steps done already. Once you’re logged in we’re ready to start playing around with Buckets!

Hands-on With Xet and Buckets

Now you should have a directory with the data file that looks like this:

Screenshot 2026-07-16 at 9.22.19 PM

As you can see, the file is 191M in size. You can open the file and have a look at the data; it’s NYC Taxi journey data:

Screenshot 2026-07-16 at 9.37.26 PM

Let’s create our first Bucket:

hf buckets create nyc-taxi-data --private

You should get some output like this:

Screenshot 2026-07-16 at 9.25.07 PM

You can guess what the “ - - private” flag does… if you want this Bucket to be public, don’t include this flag! If you go to your Huggingface page (mine is https://huggingface.co/prpatel for example, substitute your username for prpatel), you should see your newly created Bucket:

Screenshot 2026-07-16 at 9.26.23 PM

And you’ll see it’s 0 Bytes in size as we haven’t put anything into there yet. Let’s do that now by copying the NYC.csv file to it:

hf buckets cp NYC.csv hf://buckets/prpatel/nyc-taxi-data

Screenshot 2026-07-16 at 9.30.45 PM

On my fast internet connection, it took about 10 seconds to copy this ~200M file. If you open the Bucket in your web browser from your Huggingface homepage, ( for me it’s : https://huggingface.co/buckets/prpatel/nyc-taxi-data) you will see it’s now there:

Screenshot 2026-07-16 at 9.33.28 PM

Ok, so we haven’t done anything very interesting yet, so let’s see the magic of Xet in action now. Use this command to append 100 lines to the file…. since this is a demo to help you learn about Xet and Buckets, we’ll just copy a few lines from the top of the file to the bottom. You can do this by grabbing the first 100 or so rows in the CSV and copying them to the bottom, but it’s a large file and will take some time to open in an editor, so run this command in the directory where the file is saved locally:

 head -n 101 NYC.csv | tail -n +2 >> NYC.csv

If we print out the file size now, we’ll see it’s bigger than before (207M vs 191M):

Screenshot 2026-07-16 at 9.44.21 PM

Now here’s where Xet comes into play. Run the copy command exactly like before. The tool will recognize that the file already exists and Xet will use content-defined chunking (CDC) to only upload the modified part of the file, these chunks are variable sized but usually around 64KB in size. Again, here’s the command to run:

hf buckets cp NYC.csv hf://buckets/prpatel/nyc-taxi-data

Don’t forget to put in your username instead of mine! You will notice that the operation was much faster this time… for me, it took about 10 seconds to upload the full file initially, on this update, it took less than a second!

Screenshot 2026-07-16 at 9.48.27 PM

Again, Huggingface storage buckets are Xet-enabled by default. The client automatically performs chunk-level deduplication. It evaluates the file in 64kB chunks and uploads only the pieces that have changed or do not already exist on the Hub. Note the nuance here: this checking and chunking is done ON THE CLIENT, not on the server. Hence, you must have a full copy of your data file(s) locally to work on them and upload them.

Text vs Binary files and sync

As you already know, the CSV file format is text and uncompressed - I always convert them into a more efficient format like Parquet. While CSV files are row-based, Parquet files are Column-based, and Parquet files store the data in binary format (whereas CSV files store them in text format). There’s alot more to Parquet files than this, and if you’d like a deep dive into them, please leave a comment and I’ll write a separate article about them!

To convert the file from CSV to Parquet, we’ll need some tools. We’ll use Python to write this little utility, and we’ll install them using pip (please feel free to use whatever you prefer, such as uv)

pip install pyarrow pandas

And here’s our convert_to_parquet.py file:

import pandas as pd

# 1. Load the CSV into a DataFrame

df = pd.read_csv("NYC.csv")

# 2. (Optional) Do your data cleaning here...

# 3. Save directly to Parquet

df.to_parquet("NYC.parquet", engine="pyarrow")

print("Conversion complete!")

Now let’s run it:

python convert_to_parquet.py

and we’ll see a new file in that folder - notice how much smaller it is compared to the CSV file (again, this is because Parquet is a binary storage format)

Screenshot 2026-07-16 at 10.03.49 PM

Now let’s copy this new data file to Huggingface Buckets, but instead of using the cp command, we’ll use the sync command:

hf buckets sync . hf://buckets/prpatel/nyc-taxi-data

Screenshot 2026-07-16 at 10.05.36 PM

Notice a few things here:

  • it detected that our NYC.csv file did NOT need to be synced (that’s Xet in action again)
  • it also sync’d our utility file

Wrap up

Congratulations, you’ve learned the basics of Xet technology and Huggingface Buckets! There’s some important things we should discuss before you start using Buckets in your projects. First and most important, Buckets are a MUTABLE storage mechanism - there is NO VERSIONING. When you cp or sync something that exists, it will overwrite what was there previously! Once you’ve completed whatever work you are doing - say you are transforming a large dataset to train a model - you will want to version it. This can be done using a Huggingface Dataset Repository. You can very quickly copy from Buckets to Dataset repo’s - it is a server-to-server transfer so it is super fast! We’ll cover that in an upcoming article.

One other important thing to note is that the hf commands we ran can also be executed using the Huggingface Python libraries, which means you can programmatically do the things we did from the command-line. This is important for automation, scripting, re-running data cleaning & transformations, updating training, etc. We’ll also cover this in a future article.

Thanks for reading and leave your comments below!

Community

Sign up or log in to comment