AI Service Providers and Models

IMatch supports a range of AI service providers, both cloud-based commercial offerings and free locally installed AI services.

Quick Links

Click the AI service you plan to use for detailed information.

Running AI locally on your PC (no additional cost, no privacy issues)

Running Models Locally with Ollama or LM Studio

As explained below in more detail, the free software Ollama enables you to run powerful AI models on your PC at no cost (except for energy) and without privacy issues. Your images never leave your computer.

However, AI is very demanding when it comes to hardware, especially your graphics card (GPU). To run any of the modern models locally, you'll need a beefy graphics card with at least 8 GB VRAM on board; preferably, 16 GB VRAM.

How Much Memory Does Your Graphics Card Have?

Open Windows Task Manager with Shift + Ctrl + ESC and switch to the Performance tab. Click the GPU on the left to see how much memory your graphics card has:

How to tell the amount of VRAM on your graphics card.

If this counter shows less than 8 GB, it is not really up to the task of running AI locally. At least not with the models and AI available today.


AI Services and Models Supported by IMatch AutoTagger

Note: Ai companies release new models at a very high frequency ("A New Week, a New Model"). The model names used in this documentation may have been declared outdated or retired when you read this help.
The AutoTagger configuration dialog always lists the currently supported models.

AutoTagger supports cloud-based AI models from companies like OpenAI, Google, and Mistral and models you can run locally using the free software Ollama or LM Studio.

Modern AI Services
These recommended services use large language models (LLMs) and prompts to produce descriptions and keywords.

OpenAI

A very good and popular service (ChatGPT) that IMatch AutoTagger can utilize to automatically add high-quality descriptions, keywords, and traits to images.

Creating an API key is thankfully quite easy and the cost to use OpenAI models is affordable, even for personal use. To get started, follow the instructions on this page.

Ask ChatGPT or Microsoft Copilot How do I create an API key for OpenAI? to get detailed instructions.

The above page allows you to sign up. Then follow the instructions to create an API key.
Currently, the recommended model is GPT-5 Mini, which offers excellent results at a very affordable price.

Open the AutoTagger configuration dialog box via Edit > Preferences > AutoTagger and select OpenAI as the AI. The model drop-down now lists all OpenAI models supported by AutoTagger. Create an API key for the model(s) you want to use.

OpenAI settings for AutoTagger.

Don't forget to configure your rate limits for OpenAI, which depend on your usage tier. See rate limits for more information.

OpenAI frequently releases new models. We keep you informed via the IMatch release notes when we add support for new models to AutoTagger.

OpenAI offers a cost-effective solution, even for personal use, and it works on any computer, without requiring a powerful graphics card or other special hardware.

You can deposit $10 into your account and disable automatic renewal. This means when your balance is depleted, OpenAI will stop processing requests, and AutoTagger will display error messages. By disabling automatic renewal, you have full control over your spending.

With just $10, you can automatically add descriptions and keywords to thousands of images.

If you already have a 'pro' ChatGPT account, note that you'll still need to create a developer account in order to get an API key for IMatch to use. OpenAI keeps these separate.

Mistral AI

Mistral is a company based in France. They are known for releasing many of their AI models for free to the open-source community, to support further enhancements in AI and to make AI technology available for local use (on our PCs).

Creating an API key is thankfully quite easy and the cost to use Mistral models is affordable, even for personal use. To get started, follow the instructions on this page to create an account.

Ask ChatGPT or Microsoft Copilot How do I create an API key for Mistral? to get detailed instructions.

When you have an account, you can subscribe to a plan. For initial testing, subscribe to the free plan. No credit card needed.

The free plan allows you to create an API key, which you can then use with AutoTagger. The free plan has (obviously) rather severe rate limits (requests that can be made per day and tokens used), but it works well enough to try out the results you can expect from Mistral.

Open the AutoTagger configuration dialog box via Edit > Preferences > AutoTagger and select Mistral AI as the AI. The model drop-down now lists all Mistral models supported by AutoTagger. Create an API key for the model(s) you want to use.

Mistral settings for AutoTagger.

Mistral offers a cost-effective solution, even for personal use. And it works on any computer, without requiring a powerful graphics card or other special hardware.

AutoTagger supports the current version of the Medium and Large models via the latest model id. When Mistral releases new model versions, AutoTagger automatically uses the new models.

Privacy

Mistral, as a European company based in France, adheres to the strict European data privacy rules and regulations. This can make Mistral the preferred choice for users—especially corporate users—who must follow the same rules or who consider privacy important.

Google Gemini

To get started, follow the instructions to create an account and get a Gemini API key. The blue button starts the process.

Getting a Google Gemini API key.

Just sign in with your Google account (or create one) and follow the instructions. Create a project and you get the API key you need to enter in IMatch AutoTagger. Ignore all the other information about "setting up your API key". AutoTagger takes care of all of that.

After getting your API key, select the Gemini AI in AutoTagger, add your API key, and you're ready to go:

Configuring AutoTagger for using the Gemini AI.

Google Gemini excels at anything related to location (places, landmarks, tourist spots, well-known buildings), animal breeds, OCR, and many other things.

If you autotag typical vacation and travel photos frequently, give the Gemini AI a test drive. Also, if you are interested in bird photography, cars, planes, boats, and similar, Gemini AI with the Flash Light model (very affordable) should produce outstanding results.

Search for Gemini prompting tips to get some ideas. The default prompts shipped with IMatch work well for a wide range of motives. But you can of course specialize your prompts to better match your preferred style and detail level.

Ollama
(Local AI)

IMatch Learning Center.
IMatch Learning Center
There is a free video tutorial showing how to install and use Ollama.

Ollama is an open-source and free project that enables you to run artificial intelligence (AI) applications locally on your PC, without relying on cloud connectivity. Since your images never leave your computer, there are no privacy issues.

With a moderately powerful graphics card (NVIDIA preferred) and at least 6 GB VRAM, Ollama can run smaller AI models like Gemma 4 4b or Qwwen 3.5 4b efficiently. For running larger models, more VRAM (12 or more GB) is required.

1. Installing Ollama

To get started, visit the official Ollama website and download the installation package for Windows. Double-click the downloaded file to begin the installation process, which requires no administrator privileges.

After installation, Ollama will appear in the Windows Taskbar and automatically start with Windows. This ensures that you can access Ollama quickly and easily.

2. Downloading a Model

As explained in the AutoTagger Prompting help topic, Ollama requires a model to function. Fortunately, many free models are available for various applications, including image analysis, coding, and chat.

For an overview of all available Ollama models, visit Ollama's Model Library page. For IMatch, only multimodal (vision-enabled) models are relevant.


Installing the Google Gemma 4 Model

This is a multimodal model released by Google in April 2026. It offers top-of-the-line vision processing, even for models reduced to 9b or 4b parameters. It supports 140 languages so you might get sensible and grammatically-correct descriptions, headlines and traits in languages other than English.

The 4b variant needs about 5 GB RAM on the graphics card for optimal performance. For the 9b model, we recommend a graphics card with 12 GB RAM.

  • Open a command prompt window by pressing Windows key + R
  • Type the word cmd and press Enter
  • Type ollama run gemma4:e4b (or ollama run gemma4:12b if your graphic card has 12 or more GB VRAM) at the command prompt and press Enter

The model will begin downloading, which can take a while, depending on your internet speed. The gemma4:e4b model is approximately 10 GB in size.

Once the download completes, you're ready to start using Ollama with IMatch AutoTagger. Continue with the AutoTagger help topic to learn how to add descriptions, keywords, and traits to your images automatically.

After the model has been downloaded, you can close the command prompt window.

AutoTagger has configurations for both the 4b and 12b variants and you can just create a new setting with the model of your choice:

Selecting the Gemma model for an AutoTagger setting.

If you find the results of Gemma 3 lacking for your type of photography, try one of the other vision-enabled models available for Ollama.



Installing the Qwen 3 Model

Qwen is a modern model released in 2026. We recommend giving it a try and see how well it works with your images.

ollama run qwen3.5:9b
for the (better) 9b version and
ollama run qwen3.5:4b
for the smaller 4b model.

Note: These are thinking/reasoning models, which take longer to produce better answers. According to the Ollama and LM Studio documentation, appending the term /no_think to your prompt reduces or eliminates thinking.


Ollama Settings in AutoTagger

If you have downloaded multiple models, you can select any of them in AutoTagger:

Ollama settings in AutoTagger.

This list shows all models AutoTagger supports, not just the models you have downloaded.

LM Studio
(Local AI)

LM Studio is an open-source application that enables you to run powerful AI models locally on your PC. It is free for private use.

Unlike Ollama, LM Studio has a comfortable user interface and is designed to be used both as an AI runner (like Ollama) and as an interactive tool to chat with the models installed in LM Studio.

LM Studio is 'better' than Ollama when you also want to use AI models outside of IMatch.

To download and install LM Studio on your PC, follow the easy-to-follow instructions on the LM Studio website.


Downloading Models

LM Studio has a comfortable model management feature. Click on the Model Search button on the left and then search for Gemma 4 or Qwen 3.5 (or whatever model is current at the time you're reading this).

Start the LM Studio Server

This is the important part. The LM Studio server makes models accessible for other applications, like IMatch.

At the bottom of the LM Studio app window, click on the Power User button:

Enabling Power User mode in LM Studio.

Now click on the Developer button on the left:

Switching to Developer Mode in LM Studio.

Enable the server with the switch (1) and then click (2) and select the Gemma 3 model you have downloaded. Use the default settings for everything.

Activating server mode and loading a model.

Using LM Studio and Models in IMatch AutoTagger

AutoTagger is already configured to support LM Studio and suitable models. Not all models you can use with LM Studio are multimodal (vision-enabled). For AutoTagger, only vision-enabled models are usable, such as Gemma 3, LLaVA, or Llama Vision.

However, you can install any other model offered by LM Studio and use it for chatting, spelling and grammar checks, text generation, SEO, math, etc. The LM community makes suggestions regarding which models to use for each purpose. Things are changing rapidly, and new models are released frequently.

Vision-enabled models are marked with this icon:

This icon indicates vision-enabled models.

Open Edit menu > Preferences > AutoTagger and switch the AI to LM Studio. In the Model drop-down, select the Gemma 3 model available in LM Studio. We have installed the 12B Gemma 3 model in this case:

Configuring LM Studio and Gemma 3 for AutoTagger

Check the prompt and any settings you need, then confirm the configuration with OK.

Select one or more images in a File Window and start AutoTagger with F7. Select the LM Studio AI and the setting you want to use and run it:

Running AutoTagger with LM Studio and Gemma.

Make LM Studio Run Automatically

Click on the Gear icon button at the bottom right of the LM Studio app window to open the settings. Scroll down until you see this option and enable it:

Making LM Studio run automatically.

You can now use LM Studio models in AutoTagger without starting LM Studio manually. LM Studio automatically loads the model requested by AutoTagger.


Ollama or LM Studio for Local AI?

Ollama just works mostly in the background and is the simplest solution. Use it when you only want to work with AI models within IMatch.

LM Studio allows you to chat and interact with locally installed AI models much like you would use cloud-based AI such as ChatGPT, Mistral, Copilot, or Gemini. This enables you to access AI capabilities and functions without any privacy issues or cost.

Custom AI Service
(Local or Cloud)

Expert Feature

This is AutoTagger's most flexible AI connector. It allows you to connect AutoTagger with any local or remote AI service of your choice, provided that it supports an OpenAI-compatible endpoint like .../chat/completion and structured responses. Many of the tools out there do, like for example Unsloth Desktop or llama-server. If you use such tools already, you can use them with AutoTagger too.

When you select this service type for an AutoTagger setting, you configure these additional settings:

Model

The optional name of the model to use. What to input here depends on the service you use.

For llama-server, you leave this empty. AutoTagger uses whatever model is loaded by the running llama-server instance.

Server url

The full URL for the endpoint. This usually looks something like http://127.0.0.1:8888/v1/chat/completions. In this case, the server runs on our PC on port 8888 and the name of the OpenAI-compatible endpoint is v1/chat/completions. This example works with Unsloth Desktop with the default settings. For llama-server the default URL is http://127.0.0.1:11234/v1/chat/completions

Check the documentation of the service you're using for the correct address and endpoint name.

Your API key

The optional API key to use. If you have to provide a key depends on the service you're using. llma-server does not require a key, Unsloth Desktop requires you to provide the key you have created.

Threads

The number of parallel threads to use. The default value of 1 means one thread only. This is a safe default and a good starting point.

Setting threads to 0 means auto. In this mode, IMatch decides how many parallel threads to use when accessing the service. Which may work for your configuration, or not.

Set this to 1 to use only one thread. Then test. For local AI, use Windows Task Manager to check on GPU, CPU and memory utilization. Then increase this value and test again, until your PC runs out of resources or you get errors from your AI service.

For local AI, try to keep the GPU/CPU well utilized, without any prolonged idle times. More threads usually means more throughput, when all components and services can handle it.

For llama.-server, set threads to the same value as the --parallel command line parameter. If you use Unsloth Desktop, set it to the same value you have configured for the slots option in Unsloth. For other systems, please refer to their documentation.

When To Use This?

If you want to run AI models with a service not directly supported by IMatch, like Unsloth Deskop or similar, this is the way to make it work with AutoTagger.

Usage Tips for llama-server

The main goal of llama.cpp is to enable running local AI with minimal setup and state-of-the-art performance on a wide range of hardware - locally and in the cloud.

llama.cpp is the core on which many AI solutions are built today. It is available on all major operating systems and supports a wide range of processors and graphics cards.

The relevant component of llama.cpp with respect to IMatch AutoTagger is llama-server. This server allows you to run any AI model supported by llama.cpp and provides an interface that IMatch can directly access via AutoTagger, similar to Ollama and LM Studio.

Using llama.cpp to run AI models is the solution with the least overhead, the best performance, and does not require any other 3rd party software. It's also something that requires some deeper computer knowledge, familiarity with command line applications etc.

Installation

Installation is simple. You download the latest release matching your operation system and graphic card from llama.cpp's GitHib repository.

For most users, the Windows x64 (Vulkan) version fits. It supports both CPU and a wide range of graphic cards. Download the ZIP file, extract the contents into a folder, and you're good to go.

You can install it also via the winget command in Windows:
winget install llama.cpp
Windows downloads and installs it, and adds it to the PATH so you can run it from everywhere. And Winget will keep llama.cpp up-to-date automatically. With winget uninstall llama.cpp you can uninstall it cleanly.

For users with a NVIDIA Graphic Card

If your PC is equipped with a NVIDIA graphics card and you want the best possible performance, use one of the Windows x64 (CUDA 12) - CUDA 12.4 DLLs (older NVIDIA graphic cards) or Windows x64 (CUDA 13) - CUDA 13.3 DLLs (more recent NVIDIA cards) versions. Download both ZIP files (llama and the CUDA DLLs) , extract them into a folder, and you're good to go.
There are also versions from AMD and Intel to get the best performance out of their graphic cards.

Models

llama-server can use models in GGUF format you have already downloaded, or download models on first use from Hugging Face using the -hf command line argument. See documentation for more information.

LM Studio uses models in GGUF format. If you have downloaded models using LM Studio, you can use the same models with llama-server.

Running llama-server

This is both the easy and, initially, complex part:

  1. Open a command prompt window
  2. cd into the folder where you unpacked the llama.cpp binaries
  3. Assuming you have downloaded your model already and placed it into C:\data\ai-models;
    llama-server -m "C:\data\ai-models\gemma-4-12B-it-Q4_K_M.gguf" --mmproj "C:\data\ai-models\mmproj-gemma-4-12B-it-BF16.gguf" --port 11234 -c 8192 -ngl 99 -t 8 --parallel 2 --temp 1.0 --top-p 0.95 --top-k 99 --flash-attn on --threads-draft 4 --reasoning-format none --reasoning off
    This starts llama-server with the Gemma 4 12b model. It works the same with other models. Note that vision-enabled models always use a model and a so-called projector model (with the --mmproj parameter).

    Set the -t parameter to the number of physical processors in your computer.
    The --parallel n controls how many parallel requests llama-server supports. Experiment with this value (the default is 4), and set the threads option in the AutoTagger settings to the same value you have used for --parallel. This tells AutoTagger how many parallel requests it is allowed to make.

    Finding a good setting for -parallel for your CPU and GPU can make a big performance difference. Start with --parallel 2 and 2 threads in the AutoTagger settings. Run AutoTagger for a set of 20 or 30 files and measure how long it takes to process them. Increase --parallel and threads and repeat the test. If it is faster, increase a again. If it is slower or AutoTagger reports errors, decrease both values again. Keep an eye on the Performance tab in Windows Task Manager to see how the CPU and GPU is utilized.

This screen shot shows standard settings for using a running instance of llama-server from AutoTagger. AutoTagger will use whatever model you have loaded in llama-server. Note that the port number (11234) may be different for your installation.

Settings for llama-server in AutoTagger

This is a good starting point. All llama-server parameters are explained in their documentation and you can always ask an AI for details.

If you load a different model into llama-server while IMatch is running, you need to close and re-open the database. IMatch retrieves and caches the model list only once, during the first request for performance reasons.

Usage Tips for Unsloth Desktop

Unsloth Desktop is a free, open-source app for running and training AI models on your own local hardware. The company behind Unsloth is also known for creating high-quality distilled / quantized models from open weight models and making them available on Hugging Face.

To install Unsloth Desktop, follow the instructions on their website. To use Unsloth like Ollama or LM Studio from AutoTagger, you run it as an api endpoint, see https://unsloth.ai/docs/basics/api. Unsloth then runs as a server, and AutoTagger can connect to it as it would connect to Ollama or LM SAtudio.

One advantage of Unsloth Desktop over Ollama or LM Studio is that it can also perform tasks like local training, creating images, video, audio and much more.

Another advantage is that Unsloth Desktop is aware of the latest models Unsloth provides, an they often work 'better' or 'faster' than other models available on Hugging Face.

This screen shot shows some typical settings for using a running Unsloth Desktop instance from AutoTagger. Note that the port number (8888) may be different for your installation.

Settings for Unsloth Desktop in AutoTagger

Unsloth has an option that allows to specify the model when making requests. If this is enabled, you can specify the model name in your AutoTagger settings as shown above.

Enabling the option to switch models per API request.

Thinking Models

Thinking models like Google Gemma 3, Qwen 3.6 etc. may produce better results - but they also take a lot more to respond.

In our experiments, we did not notice much of a difference between the quality of descriptions and keywords produced in thinking mode or in non-thinking mode. While thinking mode can be very helpful to solve complex programming or math tasks, it does not as much for simpler tasks like describing images.

Since Unsloth Desktop uses llama-server internally and allows to provide additional llama-server parameters in the model settings, you can disable thinking by providing these twp parameters:

--reasoning-format none --reasoning off

Providing extra parameters to llama-server in Unsloth Desktop

You may also experiment, e.g.:

--reasoning-format none --reasoning auto --reasoning-effort minimal

where --reasoning-effort supports minimal, low, medium, high, xhigh or max (default: default)

As so often, what works best for you will require some testing with various settings and checking the returned descriptions, keywords, and traits.

In our tests on a NVIDIA RTX5080, disabling thinking reduced the response time from about 5 seconds to 1 second per image, for a complex prompt that returns a description, hierarchical keywords and a headline.

See the llama-server documentation for additional parameters and their meaning.

Things change fast in AI. We wrote this in August 2026. At the time you read this, other parameters or tokens you can include in your prompt may be more effective. Always refer to the current documentation of Unsloth Desktop and llama-server.



A Note About Thinking/Reasoning Models

The latest generation of models like OpenAI's GPT-5, Qwen 3, Gemini 3.* are so-called reasoning models. They "think" before they answer, supposedly producing better results that way.

The thinking step is designed to help with complex problems like math, physics, and science. Not necessarily for improving descriptions and keywords, which is the use case for IMatch AutoTagger.

Before you use these (usually extra expensive and slow) models, run some tests to see if the longer response time and cost is worth it, giving you better descriptions, keywords, and traits for your images. This may be the case if you want to get exact species or landmark identification, chart analysis, and similar tasks.

For creating descriptions and keywords for 'regular' images, often the cheaper and faster non-thinking or "Small/Nano/Mini/..." models will work just fine.

According to the Ollama and LM Studio documentation, you can disable thinking by adding /no_think or /nothink to the beginning or end of your prompt or system prompt. This should reduce or even skip the thinking phase. See the documentation of your preferred cloud AI to learn how to disable or reduce thinking.

Some Notes about Pricing of Cloud Models

AI companies bill you per request and/or the number of tokens used. Pricing sometimes varies, depending on the model you use (more powerful models are more expensive).

Check the website of the AI vendor you use for pricing information.

All AI vendors provide billing information (requests made and tokens used in the current billing period) on their website. Keep an eye on that when you are using a subscription model with automatic renewal.

Pay per Token

The AI vendor calculates the cost per request based on the number of tokens used for the prompt (input tokens) and the response (output tokens).

A good average is about 75 tokens per 100 characters in English text.

For example, the prompt "Describe this image in the style of a news headline" sent by AutoTagger to OpenAI produces 10 tokens (GPT-4o-mini model). The response "A close-up portrait of a majestic snowy owl sitting on a perch in a dimly lit studio, staring directly into the camera." has 26 tokens.

As of June 2026, OpenAI charges US$0.25 for one million input tokens and US$1.25 for one million output tokens for the GPT 5.4 nano model. Plus a small fee per request.

These affordable rates mean that you can process many thousands of images for a few dollars. And save days or even weeks of work.

Billing metrics will vary quite a bit. Because your prompts may be longer or you may be using traits (more input tokens) or the AI returns more elaborate responses (more output tokens). Use the billing page of the AI vendor to keep track of the actual cost accumulated over the billing period.

AutoTagger lists the number of requests made and tokens used since the last reset. However, there may be differences in how the AI service counts these metrics. Always use the info displayed in your dashboard after logging into your service to see usage and pricing data.