Skip to main content

Hugging Face Integration

Hugging Face Integration

Hugging Face Integration

Access thousands of open-source models — LLMs, embedding models, vision models, and more — and deploy them in production-grade inference pipelines on your own infrastructure.

Learn MoreHuggingface

Open-Source AI at Production Scale

Hugging Face is the backbone of the open-source AI ecosystem. With over 500,000 models, datasets, and tools, it’s the starting point for most custom model work. Tensorplay specializes in taking Hugging Face models beyond the notebook and into production.

What we handle for you:

  • Model selection and benchmarking across the Hugging Face Hub for your specific task
  • Fine-tuning with PEFT (LoRA, QLoRA) using the Transformers and TRL libraries
  • Quantization with bitsandbytes for memory-efficient deployment
  • High-throughput inference deployment using vLLM, Text Generation Inference (TGI), or NVIDIA Triton
  • Embedding pipelines using Sentence Transformers at scale
  • CI/CD for model training, evaluation, and deployment using Hugging Face’s model registry

Whether you’re fine-tuning Llama 3, deploying a custom reranker, or building a state-of-the-art RAG pipeline with open-source embeddings, we know this ecosystem inside and out.