Local AI runs model inference on your own device. After downloading software and models, supported local workflows can work offline; cloud features and extensions may still send data. RAM is system memory, VRAM is GPU memory. CPUs can run some models more slowly; a compatible GPU can accelerate supported workloads. Larger models require more memory. Quantization uses fewer bits to reduce size, trading some accuracy for easier running. Start small, check model licenses and measure on your actual machine.
Open-source core + optional services
Ollama
Run compatible language models locally; explicitly distinguish local models from optional cloud models.
Platforms
Windows · macOS · Linux
Guidance & sources
Start with a small local model and check whether cloud functionality is enabled.
Disable optional cloud features when you require a local-only workflow. Model licenses are separate.
Local AI runs model inference on your own device. After downloading software and models, supported local workflows can work offline; cloud features and extensions may still send data. RAM is system memory, VRAM is GPU memory. CPUs can run some models more slowly; a compatible GPU can accelerate supported workloads. Larger models require more memory. Quantization uses fewer bits to reduce size, trading some accuracy for easier running. Start small, check model licenses and measure on your actual machine.
Explore a graphical workflow for downloading and running compatible local language models.
Platforms
Windows · macOS · Linux
Guidance & sources
Check the model, runtime compatibility and your device memory before a large download.
Edition and requirements can change; check the official source before installing.
Local AI runs model inference on your own device. After downloading software and models, supported local workflows can work offline; cloud features and extensions may still send data. RAM is system memory, VRAM is GPU memory. CPUs can run some models more slowly; a compatible GPU can accelerate supported workloads. Larger models require more memory. Quantization uses fewer bits to reduce size, trading some accuracy for easier running. Start small, check model licenses and measure on your actual machine.
Build visual AI workflows using nodes; expect more setup than a one-button image service.
Platforms
Windows · macOS · Linux
Guidance & sources
Begin with a documented workflow; custom nodes and models need their own review.
Edition and requirements can change; check the official source before installing.
Local AI runs model inference on your own device. After downloading software and models, supported local workflows can work offline; cloud features and extensions may still send data. RAM is system memory, VRAM is GPU memory. CPUs can run some models more slowly; a compatible GPU can accelerate supported workloads. Larger models require more memory. Quantization uses fewer bits to reduce size, trading some accuracy for easier running. Start small, check model licenses and measure on your actual machine.