Forging high-quality, niche vision datasets — and the open-source tool to build them.
KeenForgeAI is an open-source organization focused on re-annotating and curating small, high-quality niche datasets for computer vision, and on building KeenForge — a local-first annotation & training desktop tool that makes this kind of careful work practical.
🔗 GitHub: https://github.com/KeenForgeAI/KeenForge · Datasets: https://huggingface.co/KeenForgeAI
CVPR 2026 Demo Track · Formerly AutoLabel Pro
KeenForge is a desktop application that lets you train your own object-detection model without writing a single line of code. Import images, draw a few boxes, and the AI trains itself in the background. Your data never leaves your computer.
| KeenForge | |
|---|---|
| Deployment | Double-click an .exe (Windows) |
| Privacy | 100% local & offline — images are never uploaded |
| AI assistance | Built-in YOLO-World zero-shot detection |
| Auto-training | Triggers automatically after ~15 labels |
| Model ownership | You own your model |
| Price | Free & open source (MIT) |
Core features
"welding defect", "safety helmet", "platelet")
and KeenForge finds it immediately. No pre-training required.🎬 Watch the demo · ⭐ Star on GitHub
Most public datasets are large but noisy. Smaller well-annotated datasets in niche domains are often the ones that actually move research forward — yet they are the hardest to find.
We take public niche datasets — industrial / manufacturing, medical & microscopy, agriculture, safety, and more — and:
| Dataset | Domain | Size | Description |
|---|---|---|---|
GC10-DET-corrected |
🏭 Industrial (steel) | 2,280 images · 3,542 boxes | Hot-rolled steel strip surface defects, 10 classes (Pascal VOC) — cleaned version of GC10-DET |
TXL-PBC-corrected |
🩸 Medical (hematology) | 1,256 images · 18,098 boxes | Peripheral blood cell detection (WBC / RBC / Platelets) — corrected version of TXL-PBC |
raccoon-corrected |
🦝 Wildlife | 193 images · 211 boxes | Fully re-annotated version of the classic Raccoon detection dataset |
More coming — PCB, agriculture, safety, …
Only high-quality data can push visual AI forward.
Architectures are converging and compute keeps getting cheaper — but a model is only ever as good as the labels it learns from. We optimise for annotation quality, provenance and reproducibility, not raw size.
KeenForgeAI is a community effort. You can help by:
The best vision datasets will come from many small communities working carefully, not from a few giant scrapes. Come build with us.
KeenForgeAI 是一个开源组织,专注于「高质量小众视觉数据集」的重新标注与整理,并开发配套的本地化标注训练工具 KeenForge。
CVPR 2026 Demo Track · 前身为 AutoLabel Pro
KeenForge 是一款桌面应用,无需写一行代码就能训练你自己的目标检测模型:导入图片 → 画几个框 → AI 在后台自动训练。数据全程不离开你的电脑。
公开数据集往往「大而糙」——框松、漏标、错标、重复。真正推动研究的,常常是那些小众领域里标注精良 的小数据集,但它们恰恰最难找。
我们把公开的小众数据集(工业制造、医疗显微、农业、安防等)重新整理:
已发布数据集:🏭 GC10-DET-corrected(工业 — 钢板表面,2,280 张 / 3,542 框)、🩸 TXL-PBC-corrected(医疗 — 血液学,1,256 张 / 18,098 框)、🦝 raccoon-corrected(野生动物,193 张 / 211 框)
只有高质量的数据集,才能推动视觉 AI 的发展。
模型架构在趋同、算力越来越便宜,但模型的上限永远由标注质量决定。我们死磕的是标注质量、 可溯源性和可复现性,而不是数据量。
我们相信,未来最好的视觉数据集会来自许多小社区认真细致的工作,而不是少数几个巨大的爬取。 欢迎一起共建 KeenForgeAI。
🔗 GitHub: https://github.com/KeenForgeAI/KeenForge | Hugging Face: https://huggingface.co/KeenForgeAI