distributed-training
27 个项目 · ⭐ 175.5kLearn how to develop, deploy and iterate on production-grade ML applications.
The largest collection of PyTorch image encoders / backbones. Including train, eval, inference, export scripts, and pretrained weights -- ResNet, ResNeXT, EfficientNet, NFNet, Vision Transformer (ViT), MobileNetV4, MobileNet-V3 & V2, RegNet, DPN, CSPNet, Swin Transformer, MaxViT, CoAtNet, ConvNeXt, and more
Easy-to-use and powerful LLM and SLM library with awesome model zoo.
Fengshenbang-LM(封神榜大模型)是IDEA研究院认知计算与自然语言研究中心主导的大模型开源体系,成为中文AIGC和认知智能的基础设施。
FEDML - The unified and scalable ML library for large-scale distributed training, model serving, and federated learning. FEDML Launch, a cross-cloud scheduler, further enables running any AI jobs on any GPU cloud or on-premise cluster. Built on this library, TensorOpera AI (https://TensorOpera.ai) is your generative AI platform at scale.
A high performance and generic framework for distributed DNN training
Fast and flexible AutoML with learning guarantees.
Training and serving large-scale neural networks with auto parallelization.
Decentralized deep learning in PyTorch. Built to train models on thousands of volunteers across the world.
An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
:dart: Gradient Accumulation for TensorFlow 2
FTPipe and related pipeline model parallelism research.
共 27 条 · 第 1 / 2 页