Does this use table buckets?
HN user
chrislusf
I work on SeaweedFS. Let me know if see any bugs or just create a github issue.
If less components is desired, use SeaweedFS, which supports S3 table buckets and Iceberg catalog and maintenance. Basically storing Iceberg tables data and metadata.
I work on SeaweedFS.
Just download the single binary, for most platforms, and run "weed mini -dir=your_data_directory", with all the configuration optimized.
I work on SeaweedFS.
I am trying to support AWS S3 APIs as complete as possible.
Recently added support for Table Bucket, besides myriads of details, such as policies, STS, IAM, OIDC, WORM, lock and versioning, governance, etc.
(I work on SeaweedFS.)
Haha, you used Claude to find the Clause code.
I used Claude to generate a lot of admin UI pages, saved a lot of time. The core storage engine part I dare not using AI, same as you.
I work on SeaweedFS since 2011, and full time since 2025.
SeaweedFS was started as a learning project and evolves along the way, getting ideas from papers for Facebook Haystack, Google Colossus, Facebook Tectonics. With its distributed append-only storage, it naturally fits object store. Sorry to see MinIO went away. SeaweedFS learned a lot from it. Some S3 interface code was copied from MinIO when it was still Apache 2.0 License. AWS S3 APIs are fairly complicated. I am trying to replicate as much as possible.
Some recent developments:
* Run "weed mini -dir=xxx", it will just work. Nothing else to setup.
* Added Table Bucket and Iceberg Catalog.
* Added admin UI
I work on SeaweedFS. It is not backed by any greedy VC. So no urgency to make a large profit from the open source community.
It is in the parent comment.
Not correct. The files are chunked into smaller pieces and spread to all volume servers.
Disclaim: I work on SeaweedFS.
Why skipping SeaweedFS? It rank #1 on all benchmarks, and has a lot of features.
I work on SeaweedFS. So very biased. :)
Just run "weed sever -s3 -dir=..." to have an object store.
I work on SeaweedFS. It has support for these if conditions, and a lot more.
Thank! There is an admin UI already. AI coding makes this fairly easy.
This is Chris and I am the creator of SeaweedFS. I am starting to work full time on SeaweedFS now. Just create issues on SeaweedFS if any.
Recently SeaweedFS is moving fast and added a lot more features, such as: * Server Side Encryption: SSE-S3, SSE-KMS, SSE-C * Object Versioning * Object Lock & Retention * IAM integration * a lot of integration tests
Also, SeaweedFS performance is the best in almost all categories in a user's test https://www.repoflow.io/blog/benchmarking-self-hosted-s3-com... And after that, there is a recent architectural change that increases performance even more, with write latency reduced by 30%.
I really don't understand why you aren't eager to explain the differences and what problems are being solved.
Sorry, everybody has different background of knowledge. Hard to understand where the question comes from. I think https://www.usenix.org/system/files/fast21-pan.pdf may be helpful here.
Why does a user need that? Filesystems already break up files into blocks / sectors. Why wouldn't a user just deal with files and let the filesystem handle it?
A blob has its own storage, which can be replicated to other hosts in case current host is not available. It can scale up independently of the file metadata.
The blob storage is what SeaweedFS built on. All blob access has O(1) network and disk operation.
Files and S3 are higher layers above the blob storage. They require metadata to manage to the blobs, and other metadata for directories, S3 access, etc.
These metadata usually sit together with the disks containing the files. But in highly scalable systems, the metadata has dedicated stores, e.g., Google's Colossus, Facebook's Techtonics, etc. SeaweedFS file system layer is built as a web application of managing the metadata of blobs.
Actually SeaweedFS file system implementation is just one way to manage the metadata. There are other possible variations, depending on requirements.
There are a couple of slides on the SeaweedFS github README page. You may get more details there.
A large file can be chunked into blobs.
Sorry it was not so clear. Previously fallocate just allocate disk space for a local server. Now SeaweeedFS can allocate a blob on a remote storage.
How is that not mmap?
The allocated storage is append only. For updates, just allocate another blob. The deleted blobs would be garbage collected later. So it is not really mmap.
Also what is the difference between a file, an object, a blob, a filesystem and an object store?
The answer would be too long to fit here. Maybe chatgpt can help. :)
Is all this just files indexed with sql?
Sort of yes.
Should not be a problem.
One similar use case used Cassandra as SeaweedFS filer store, and created thousands of files per second in a temp folder, and moved the files to a final folder. It caused a lot of tombstones for the updates in Cassandra.
Later, they changed to use Redis for the temp folder, and keep Cassandra for other folders. Everything has been very smooth since then.
Thanks for sharing! I work on SeaweedFS.
SeaweedFS is built on top of a blob storage based on Facebook's Haystack paper. The features are not fully developed yet, but what makes it different is a new way of programming for the cloud era.
When needing some storage, just fallocate some space to write to, and a file_id is returned. Use the file_id similar to a pointer to a memory block.
There will be more features built on top of it. File system and Object store are just a couple of them. Need more help on this.
I work on SeaweedFS. Ping me if you need some help to understand why SeaweedFS is much faster.
It is still growing. This is not the case since many months ago.
Also, welcome to make a PR.
Another research suggests rubbing eyes also increase risk for Alzheimers.
The virus getting into the eyes will slowly eat the brain.
Can someone please critique on the difference between SeaweedFS and Colossus? https://github.com/seaweedfs/seaweedfs
Need some advice on how to make SeaweedFS more scalable.
Thanks for the detailed clarification! I am too deep into the SeaweedFS low level details and am all ears on how to make it simpler to use. SeaweedFS has weekly releases and is constantly evolving.
Depending on your case, you may need to add more filers. UCSD has a setup that uses about 10 filers to achieve 1.5 billion iops. https://twitter.com/SeaweedFS/status/1549890262633107456 There are many AI/ML users switching from MinIO or CEPH to SeaweedFS, especially with lots of images/text/audio files to process.
I found MinIO benchmark results is really, well, "marketing". MinIO is basically just an S3 API layer on top of the local disks. Any object is mapped to at least 2 files on disk, one for metadata and one is the object itself.
SeaweedFS author here. Thanks for your candid answer. You do not need to use multiple SeaweedFS components. Just download the binary and run "weed server -s3".
There are many other components, but you do not really need to use them. This default mode should be good enough for most cases. I saw many times people try to optimize too early, but often unnecessary, and sometimes in the wrong way.
I would like to know what kind of setup you are running. It should beat most other options if the use case needs lots of small files, e.g. millions or billions of files. If just small use case, e.g. a few personal files, it would be an overkill.
Another aspect is how to increase capacity for existing clusters. It should be most simple for SeaweedFS, just start one more volume server. And it will linearly increase the throughput.
Besides storage cost, S3 API access cost can also be high if frequently accessed. And latency is unpredictable.
You can use SeaweedFS Remote Object Store Gateway to cache S3 (or any S3 API compatible vendors) to local servers, and access them at local network speed, and asynchronously sync back to S3.
https://github.com/chrislusf/seaweedfs/wiki/Gateway-to-Remot...