Skip to main content
Ctrl+K
Embedl Docs Embedl Docs
  • Deploy
  • Hub
  • Models
  • Embedl
  • GitHub
  • HuggingFace
  • Deploy
  • Hub
  • Models
  • Embedl
  • GitHub
  • HuggingFace
  • Embedl Deploy
  • Installation
  • Quickstart
  • User Guide
    • Optimization Pipeline
    • Backends
    • Graph Conversions
    • Operator Fusions
    • Quantization
    • Custom Patterns
  • Tutorials
    • TensorRT inference helpers
    • Faster ResNet deployment
    • Deploying vision models
    • SAM3 deployment
    • SAM 3D Body deployment
    • Model deployment and tracking with Embedl Hub
  • API Documentation
    • embedl_deploy.axelera package
      • embedl_deploy.axelera.modules package
      • embedl_deploy.axelera.patterns package
    • embedl_deploy.backend package
    • embedl_deploy.lattice package
      • embedl_deploy.lattice.modules package
      • embedl_deploy.lattice.patterns package
    • embedl_deploy.quantize package
    • embedl_deploy.tensorrt package
      • embedl_deploy.tensorrt.modules package
      • embedl_deploy.tensorrt.patterns package
    • embedl_deploy.version package
      • embedl_deploy.version.public module
    • embedl_deploy.vitis package
      • embedl_deploy.vitis.modules package
      • embedl_deploy.vitis.patterns package
  • User Guide

User Guide#

This guide covers the full Embedl Deploy optimization pipeline — from graph conversions through operator fusions to INT8 quantization — with working examples on ResNet50, ConvNeXt, and Vision Transformer (ViT).

  • Optimization Pipeline
    • Pipeline stages
    • One-shot API
    • Plan-based API
    • Pattern priority
    • Pattern groups
    • Verifying numerical equivalence
    • ONNX export and compilation
  • Backends
    • Built-in backends
    • Selecting a backend
    • API reference
  • Graph Conversions
    • Built-in TensorRT conversions
    • When conversions matter
    • Running conversions only
  • Operator Fusions
    • Convolution fusions
    • Linear fusions
    • Attention fusions
    • Pooling fusions
    • Fusion summary by architecture
    • Running fusions only
  • Quantization
    • Quantization pipeline
    • Step 1: Transform and fuse
    • Step 2: Quantize
    • QDQ stub placement
    • Why pattern-aware QDQ matters
    • Full example: ResNet50 INT8 PTQ
    • SmoothQuant
    • Quantization-Aware Training (QAT)
  • Custom Patterns
    • Why customize?
    • Writing a custom pattern
    • Building a custom pattern list
    • Using the custom pattern list

previous

Quickstart

next

Optimization Pipeline

Show Source
HuggingFace

© 2026 Embedl All rights reserved