Hugging Face local models¶
Local execution is the Hugging Face path. annotate builds ChatHuggingFace around HuggingFacePipeline.from_model_id pointed at a directory on disk. Transformers reads that directory (a single model.safetensors or sharded model-00001-of-0000N.safetensors plus model.safetensors.index.json). The rest of the app still talks to the same chat session the other providers use.
HF_BACKEND=remote keeps HuggingFaceEndpoint for hosted inference. endpoint is an older name for the same mode. Remote mode needs HF_TOKEN or HUGGINGFACEHUB_API_TOKEN. Local mode does not.
Install¶
Core extraction does not install Transformers, Accelerate, bitsandbytes, or PyTorch.
Use cuda130.txt or cuda129.txt when the driver reports CUDA 13.0/13.1 or 12.9. Otherwise use cpu.txt, macos.txt, gpu.txt, or xpu.txt. python setup.py reads nvidia-smi and installs the matching file. The three CUDA files include bitsandbytes. The others do not.
Configure¶
Copy the keys you need from .env.example.
export LLM_ENABLED=true
export LLM_PROVIDER=huggingface
export LLM_TASKS=screen_classification
export HF_BACKEND=local
export HF_LOCAL_MODEL_PATH=/models/Qwen3-0.6B
export HF_QUANTIZATION=none
export HF_OFFLINE=true
figma-extractor annotate --dir ./out
HF_LOCAL_MODEL_PATH is the directory. LLM_MODEL is not a Hub repo id in local mode. If LLM_MODEL itself is an existing directory, it is accepted when the path variable is unset.
| Choice | Effect |
|---|---|
HF_QUANTIZATION=none |
Full precision. Default |
HF_QUANTIZATION=4bit |
BitsAndBytes NF4. Lower memory, more numeric error |
HF_QUANTIZATION=8bit |
BitsAndBytes 8-bit. More memory than 4-bit, closer to full precision |
HF_DTYPE=auto |
float32 on CPU. On CUDA, bfloat16 only when the GPU implements it in hardware, otherwise float16. Emulated bfloat16 is not selected. Local chat turns off template reasoning (enable_thinking=false) so the generation budget is the answer rather than a <think> block. |
HF_DEVICE=cpu |
Force CPU. device_map=auto is not applied |
HF_DEVICE=cuda |
Require CUDA. Fails when CUDA is absent |
HF_OFFLINE=true |
Sets HF_HUB_OFFLINE and TRANSFORMERS_OFFLINE while loading, and model download refuses to run |
HF_BNB_4BIT_QUANT_TYPE, HF_BNB_4BIT_COMPUTE_DTYPE, and HF_BNB_4BIT_USE_DOUBLE_QUANT apply when 4-bit is actually active. Local loads always pass local_files_only=True. A missing weight file is an error, not a download. On some Turing GPUs, cuBLAS rejects an 8-bit matrix shape. Prompts are left-padded to a multiple of 16, with those positions masked, so the fast kernel is used. A remaining rejected shape uses the same full-precision multiplication BitsAndBytes already uses for unaligned dimensions. The loaded layers stay 8-bit. Quantized loads keep the runtime dtype, so a card without hardware bfloat16 does not cast bfloat16 weights on every layer.
The pipeline is cached in the process. A second annotate with the same path, dtype, and quantization reuses it.
Commands¶
figma-extractor model download Qwen/Qwen3-0.6B --dest /models/Qwen3-0.6B
figma-extractor model validate --path /models/Qwen3-0.6B
figma-extractor model validate --path /models/Qwen3-0.6B --load
validate without --load reads config.json and tokenizer_config.json. It reports the architecture, device, dtype, requested quantization, active quantization, and package versions. --load constructs the pipeline and, when the model exposes it, prints get_memory_footprint().
Weights belong under a models/ directory. That directory, plus *.safetensors, *.bin, *.pt, and *.pth, is gitignored.
Offline¶
Download once, then set HF_OFFLINE=true before annotate and model validate. Startup will not call the Hub. HF_HOME, TRANSFORMERS_CACHE, and HUGGINGFACE_HUB_CACHE still matter if a tool downloads into the default cache. This provider prefers HF_LOCAL_MODEL_PATH over that cache.
Deploy¶
- Install Python 3.11+.
- Create a virtual environment.
pip install -e ".[huggingface-local]".- Install
requirements/cuda132.txt,cuda130.txt, orcuda129.txtfor NVIDIA, orcpu.txt,xpu.txt,gpu.txt, ormacos.txt. - The CUDA files already include bitsandbytes for 4-bit and 8-bit.
figma-extractor model download … --dest …on a machine that may reach the Hub.- Copy the directory to the host that will run, or mount it.
- Export
HF_LOCAL_MODEL_PATH,HF_BACKEND=local, andHF_OFFLINE=true. figma-extractor devices.figma-extractor model validate --path ….figma-extractor annotate --dir ./out --llm --llm-provider huggingface --llm-task screen_classification.
Linux and Windows NVIDIA hosts use cuda132.txt (torch 2.14.1+cu132), cuda130.txt (torch 2.14.1+cu130), or cuda129.txt (Linux torch 2.13.0+cu129, Windows torch 2.8.0+cu129). macOS Apple Silicon uses macos.txt and does not install bitsandbytes. There is no Dockerfile in this repository. Mount a host directory at the path you set in HF_LOCAL_MODEL_PATH if you add a container later. Do not copy weights into an image unless you mean to.
Hardware¶
| Backend | Local full precision | 4-bit / 8-bit |
|---|---|---|
| NVIDIA CUDA 13.2, 13.0, or 12.9 | Yes, when the matching CUDA file is installed | Yes, when bitsandbytes imports and HF_DEVICE resolves to CUDA |
| CPU torch wheel | Yes, dtype float32 when HF_DTYPE=auto |
No. The run fails, or stays full precision when HF_QUANTIZATION_ON_UNSUPPORTED=fallback |
| macOS MPS, ROCm, XPU | The matching torch file loads the model without bitsandbytes | Not claimed. The same failure or fallback rule applies |
The quantization decision is unit-tested with stand-in CPU and CUDA status objects. A real weight file was not loaded in that suite. Set HF_RUN_INTEGRATION=1 and HF_LOCAL_MODEL_PATH to run tests/llm/test_huggingface_integration.py.
Active quantization in figma-extractor model validate is none unless 4-bit or 8-bit was actually selected. A fallback does not print an active 4-bit line.
What changed¶
Hosted inference used to be the default (backend=endpoint). The default is now local. Remote inference remains behind HF_BACKEND=remote. Settings are parsed once. build_local_plan records path, device, dtype, and the quantization recipe. load_pipeline reuses that plan: prepare the runtime, construct the pipeline, then adapt 8-bit prompts. model validate --load passes the plan it already built. The graph still receives a LangChainSession and does not construct Transformers itself.