Developers & AI Infrastructure

Best Computer Vision Tools

Also known as: computer vision platforms

Computer vision tools let software understand images and video: find objects, read text, count products on a shelf or spot a scratch on a production line. This list covers three kinds. Platforms (Roboflow, Ultralytics, LandingLens) help you label images, train your own model and deploy it. Ready-made cloud APIs (Amazon Rekognition, Google Cloud Vision, Azure AI Vision) recognise common things like labels, faces and text with no training. Open models and libraries (OpenCV, Meta's SAM 3) are free to run on your own hardware.

A fourth option now competes with all three: multimodal AI models such as Gemini can find and outline objects from a plain-English prompt. We ranked 10 tools on accuracy, custom training, deployment, price and developer experience. Prices are as of 25 September 2026. Note that two big clouds are pulling back: Microsoft will retire Azure's Image Analysis and Custom Vision APIs on 25 September 2028, and Google's Vertex AI Vision reaches end of life on 30 September 2026.

Quick answer

Roboflow is the best computer vision tool for most teams in 2026. It covers labeling, training, deployment and edge inference in one place, has a free plan, and private projects start at $79/month billed yearly. Pick Ultralytics YOLO to train small, fast detectors yourself (check the AGPL-3.0 licence first), SAM 3 for free open-vocabulary segmentation, and the Gemini API when a prompt like "find every dented can" is good enough and you do not want to train anything.

Top picks at a glance

Scoreboard

Scores are out of 10. The overall score is the weighted average of the criteria below.

#ToolOverallAccuracy & capabilitiesCustom trainingDeployment & edgePricing & valueDeveloper experiencePrice fromBest for
1Roboflow
Roboflow
8.98.89.59.08.09.3Free (public data); $79/month Core (billed yearly)
Free tier
Teams that need a custom detector or segmenter in production quickly
2Ultralytics YOLO
Ultralytics
8.88.89.09.58.08.8Free (AGPL-3.0); Platform Pro $29/seat/month
Free tier
Python developers training fast detectors for cameras and edge devices
3Meta SAM 3
Meta
7.99.06.57.09.07.5Free
Free tier
Precise outlines, auto-labeling and video tracking from a text prompt
4OpenCV
OpenCV (open-source project)
7.77.05.09.510.07.0Free
Free tier
Engineers building image and video pipelines with zero licence cost
5Gemini API (vision)
Google
7.68.56.06.08.59.0Free tier; then $0.75 per 1M input tokens (Gemini 3.8 Flash)
Free tier
Prototypes and low-volume jobs where a prompt can replace a trained model
6LandingLens
LandingAI
7.57.58.58.05.58.0Free (non-commercial); Enterprise custom
Free tier
Manufacturers running visual quality inspection
7Amazon Rekognition
Amazon Web Services
7.37.57.06.08.08.0$0.001 per image
Free tier
AWS teams needing labels, faces, text or moderation at scale
8Google Cloud Vision API
Google Cloud
7.07.56.55.57.58.0Free (1,000 units/month); then $1.50 per 1,000
Free tier
Reliable OCR and image labels on Google Cloud
9Clarifai
Clarifai
6.27.07.06.55.05.0Usage-based (current rates not confirmed)Existing customers running vision and language models on one platform
10Azure AI Vision
Microsoft
5.97.03.56.56.55.5Free (5,000 transactions/month); then pay per 1,000
Free tier
Existing Azure users who need time to migrate

Expert reviews

#1 · Teams that need a custom detector or segmenter in production quickly

Roboflow

by Roboflow · Freemium · Free (public data); $79/month Core (billed yearly)
8.9/10

Roboflow is the quickest way to go from a folder of photos to a working vision model. It covers the whole loop: upload and label images (with Segment Anything-based tools and auto-labeling), train a model on hosted GPUs, test it, then deploy it to a hosted API or your own hardware.

Its own model family, RF-DETR, is open source. Roboflow says it was the first real-time detector to pass 60 mAP on the COCO benchmark, the research paper was accepted at ICLR 2026, and segmentation and keypoint versions followed this year. The Inference server that runs models is also Apache 2.0 and works on NVIDIA Jetson, Raspberry Pi, CPUs and GPUs, so you are not locked into Roboflow's cloud.

Pricing is published. The free Public plan gives 15 credits a month but makes your data and models public. Core costs $79/month billed yearly ($99 monthly) for three users and private data. Credits pay for labeling, training and inference, so heavy users should budget for add-ons.

Pick it if you need a custom detector in days, not months. Skip it if your images must stay private and you have no budget; train locally with Ultralytics or RF-DETR instead.

Score breakdown

Accuracy & capabilities8.8
Custom training9.5
Deployment & edge9.0
Pricing & value8.0
Developer experience9.3

Key facts

Pricing
Free (public data); $79/month Core (billed yearly) (Public: free, 2 users, 15 credits/month, but data and models are public on Roboflow Universe. Core: $79/month billed yearly (50 credits) or $99 monthly (30 credits), 3 users, private data. Credits pay for labeling, training, deployment and inference. Enterprise: custom.)
Free option
Yes
Platforms
Web, API, Python SDK, Self-hosted, Edge (Jetson, Raspberry Pi)
Own model
RF-DETR: core models Apache 2.0; paper accepted at ICLR 2026
COCO claim
First real-time detector above 60 mAP (Roboflow)
Inference server
Apache 2.0 core; runs on Jetson, Raspberry Pi, CPU and GPU
Free plan catch
Data and models are public on Roboflow Universe

What we like

  • Labeling, training and deployment in one product
  • Apache 2.0 RF-DETR models and self-hostable Inference server
  • Runs on edge devices such as Jetson and Raspberry Pi
  • Clear published prices

Watch out for

  • Free plan makes your data and models public
  • Credits cap how much you can train and run on Core
  • Some RF-DETR Plus components use a non-Apache licence
#2 · Python developers training fast detectors for cameras and edge devices

Ultralytics YOLO

by Ultralytics · Open source · Free (AGPL-3.0); Platform Pro $29/seat/month
8.8/10

Ultralytics makes YOLO, one of the best-known families of real-time object detectors, and its newest version, YOLO26, arrived in January 2026. One Python package covers detection, instance and semantic segmentation, classification, pose, rotated boxes and depth estimation. Training on your own labeled images takes a few lines of code.

The numbers are strong for the size. Ultralytics reports 40.9 mAP on COCO for the tiny YOLO26n and 57.5 for YOLO26x, and says YOLO26n runs up to 43% faster on CPU than YOLO11n. Models export to ONNX, TensorRT, CoreML, LiteRT and OpenVINO, so they run on phones, Jetson boards and plain CPUs.

The catch is the licence. The code and the models you train are AGPL-3.0 by default. If you ship YOLO inside a commercial product without open-sourcing your own code, you need an Enterprise License, and its price is not public. The hosted Ultralytics Platform starts free with a $25 one-time credit; Pro is $29 per seat per month plus GPU time.

Pick it if you can write Python and want the shortest path to an edge-ready detector. Skip it if AGPL is a problem and you will not buy a licence; RF-DETR's core models are Apache 2.0.

Score breakdown

Accuracy & capabilities8.8
Custom training9.0
Deployment & edge9.5
Pricing & value8.0
Developer experience8.8

Key facts

Pricing
Free (AGPL-3.0); Platform Pro $29/seat/month (Code and models are free under AGPL-3.0. Ultralytics Platform: Free ($25 one-time credit, 3 concurrent trainings); Pro $29/seat/month with $30/seat of monthly credits; cloud GPUs from $0.24/hour. Enterprise License for closed-source commercial use: custom price.)
Free option
Yes
Platforms
Python, CLI, Web (Ultralytics Platform), Edge export
Latest model
YOLO26, released January 2026
COCO mAP
40.9 (YOLO26n) to 57.5 (YOLO26x), Ultralytics
CPU speed
Up to 43% faster than YOLO11n on CPU (Ultralytics)
Licence
AGPL-3.0, or paid Enterprise License

What we like

  • Seven vision tasks from one package and one API
  • Small models that run fast on CPUs and edge devices
  • Exports to every major edge runtime
  • Cheap hosted training from $0.24/GPU-hour

Watch out for

  • AGPL-3.0 by default; closed commercial use needs a paid licence
  • Enterprise License price is not published
  • You still need labeled data from somewhere
#3 · Precise outlines, auto-labeling and video tracking from a text prompt

Meta SAM 3

by Meta · Open source · Free
7.9/10

SAM 3 is Meta's third Segment Anything model, released in November 2025, and it changed what a free model can do. Earlier versions outlined whatever you clicked on. SAM 3 takes a short text prompt such as "yellow school bus", or an example image, and then finds, outlines and tracks every matching object in a photo or video. This is called open-vocabulary segmentation: there is no fixed list of classes and no training step.

Meta says SAM 3 doubled the accuracy of earlier systems on its SA-Co benchmark and takes about 30 milliseconds per image on an H200 GPU, even with more than 100 objects. SAM 3.1 (March 2026) is a drop-in update that tracks up to 16 objects in one pass and, according to Meta, doubles video throughput to 32 frames per second on one H100.

It is free to download under Meta's own SAM License, which allows commercial use with some restrictions, so read it before you ship. You need a capable GPU, and it gives you masks, not a finished app.

Pick it if you need precise outlines, fast auto-labeling or video tracking. Skip it if you need a tiny model for a phone; train a YOLO or RF-DETR model instead, using SAM 3 to label its data.

Score breakdown

Accuracy & capabilities9.0
Custom training6.5
Deployment & edge7.0
Pricing & value9.0
Developer experience7.5

Key facts

Pricing
Free (Weights and code on GitHub and Hugging Face under Meta's SAM License, which allows commercial use with some restrictions. You pay only for your own GPUs.)
Free option
Yes
Platforms
Python, Self-hosted, Web demo (Segment Anything Playground)
Released
SAM 3 November 2025; SAM 3.1 27 March 2026
Prompts
Short text phrases, example images, clicks or boxes
Speed
About 30 ms per image with 100+ objects on an H200 (Meta)
Licence
SAM License (custom, commercial use allowed)

What we like

  • Finds and outlines every object matching a text prompt
  • Tracks objects through video
  • Free weights, with commercial use allowed
  • Excellent for pre-labeling training data

Watch out for

  • Needs a data-center or strong desktop GPU for good speed
  • Custom licence, not a standard open-source one
  • Outputs masks only; you build the rest
#4 · Engineers building image and video pipelines with zero licence cost

OpenCV

by OpenCV (open-source project) · Open source · Free
7.7/10

OpenCV is the free, open-source library that many vision projects rely on for the plumbing: reading cameras and video files, resizing and colour conversion, filters, feature matching, camera calibration, 3D geometry and drawing results on screen. It works from C++, Python, Java and JavaScript, on hardware from servers to single-board computers.

OpenCV 5.0 shipped in June 2026 with a rewritten deep-learning (DNN) engine. Support for ONNX operators rose from about 22% in version 4 to over 80%, so many more modern models load directly, and the release adds building blocks for running language and vision-language models. It also dropped the old C API and now needs C++17, so older code may need porting. The licence is Apache 2.0.

What OpenCV does not do is train modern neural networks, label data or give you a dashboard. You bring a trained model (from Ultralytics, Roboflow or PyTorch), and OpenCV runs it and handles everything around it.

Pick it if you are an engineer building a pipeline and want no licence fees or vendor lock-in. Skip it if you want a no-code tool or a hosted API; choose Roboflow or a cloud API instead.

Score breakdown

Accuracy & capabilities7.0
Custom training5.0
Deployment & edge9.5
Pricing & value10.0
Developer experience7.0

Key facts

Pricing
Free (Apache 2.0. No paid tier is needed to use the library.)
Free option
Yes
Platforms
C++, Python, Java, JavaScript, Windows, macOS, Linux, Android, iOS
Latest major release
OpenCV 5.0, June 2026
ONNX coverage
Over 80% of operators in 5.0, up from about 22% in 4.x
Licence
Apache 2.0
GitHub stars
About 91k (opencv/opencv, 25 Sep 2026)

What we like

  • Free under Apache 2.0, with no usage limits
  • Runs almost anywhere, including phones and small boards
  • New DNN engine loads far more ONNX models
  • Huge community and decades of tutorials

Watch out for

  • No model training, labeling or hosting
  • Version 5 breaks some old C and C++ code
  • Needs programming skills
#5 · Prototypes and low-volume jobs where a prompt can replace a trained model

Gemini API (vision)

by Google · Usage-based · Free tier; then $0.75 per 1M input tokens (Gemini 3.8 Flash)
7.6/10

Multimodal AI models now handle jobs that once needed a custom-trained detector, and Google's Gemini API is the one we would try first, because its docs support object detection with bounding boxes and segmentation outlines out of the box. You send an image and a plain-English request, such as "find every dented can", and it returns labels, boxes or masks as structured data. It also reads text, answers questions about photos and understands video.

Google's docs recommend Gemini 3.8 Flash for image understanding. It costs $0.75 per million input tokens and $3.75 per million output tokens until 31 December 2026, then doubles, and there is a free tier. A small image counts as 258 tokens, so the input for one image costs a tiny fraction of a cent; the answer adds more.

The trade-offs are real. A general model is not tuned to your objects, so expect less consistent boxes than a detector trained on your own images. Each call also travels to Google's cloud, so it will not work offline or keep up with a 30-frames-per-second camera.

Pick it if you need a prototype today, handle many object types or process low volumes. Skip it if you need real-time, offline or pixel-exact results. See our Gemini 3.8 Flash page.

Score breakdown

Accuracy & capabilities8.5
Custom training6.0
Deployment & edge6.0
Pricing & value8.5
Developer experience9.0

Key facts

Pricing
Free tier; then $0.75 per 1M input tokens (Gemini 3.8 Flash) (Gemini 3.8 Flash: $0.75 per 1M input tokens and $3.75 per 1M output tokens until 31 December 2026, rising to $1.50 and $7.50 from 1 January 2027. A small image counts as 258 tokens; larger images are split into 768x768 tiles of 258 tokens each.)
Free option
Yes
Platforms
API, Python, JavaScript, Google AI Studio
Tasks
Captions, Q&A, object detection with boxes, segmentation masks, video
Recommended model
Gemini 3.8 Flash (Google docs)
Box format
Coordinates scaled 0 to 1000
Image cost
258 tokens per small image or tile

What we like

  • No training: describe what to find in plain English
  • Boxes and segmentation masks returned as JSON
  • Free tier and low per-image cost
  • Handles images, documents and video in one API

Watch out for

  • Cloud only, with no offline or on-device option
  • Too slow for real-time camera streams
  • Less consistent than a detector trained on your data
  • Price doubles on 1 January 2027
#6 · Manufacturers running visual quality inspection

LandingLens

by LandingAI · Enterprise · Free (non-commercial); Enterprise custom
7.5/10

LandingLens is LandingAI's platform for visual inspection in factories: finding scratches, dents, missing parts and other defects on a production line. You label images, train a model and deploy it from one guided web app, so quality engineers can own the project without a machine-learning team.

Deployment fits factory floors well. Models can run in LandingAI's cloud, on site through the LandingEdge app, or in a Docker container. The free plan is usable for trials: 1,000 credits a month, up to 3 invited users, unlimited projects and 10,000 images per project. With Fast Training, training on or running one image costs 1 credit.

The limits are commercial. The free plan's single model download is for non-commercial use only, and production use needs the Enterprise plan, which has no public price. LandingAI's main pricing page now covers only its document product, Agentic Document Extraction, which suggests where the company's attention has moved.

Pick it if you run quality inspection and want a guided tool your engineers can manage. Skip it if you need general-purpose detection or published prices; Roboflow is more open about cost.

Score breakdown

Accuracy & capabilities7.5
Custom training8.5
Deployment & edge8.0
Pricing & value5.5
Developer experience8.0

Key facts

Pricing
Free (non-commercial); Enterprise custom (Free: 1,000 credits/month (no rollover), up to 3 invited users, unlimited projects, 10,000 images per project, 1 model download for non-commercial use. Enterprise: custom price, commercial use, model downloads from 5. Fast Training costs 1 credit per image to train or run.)
Free option
Yes
Platforms
Web, LandingEdge, Docker, API
Focus
Visual inspection for manufacturing
Deployment
Cloud, LandingEdge app or Docker container
Free plan
1,000 credits/month, non-commercial only
Credit cost
1 credit per image to train or infer (Fast Training)

What we like

  • Built around factory defect inspection
  • Edge and Docker deployment for on-site lines
  • Guided workflow for non-specialists
  • Free plan for trials

Watch out for

  • Commercial use requires a custom-priced Enterprise plan
  • Narrower than general vision platforms
  • Vendor's focus has shifted toward document extraction
#7 · AWS teams needing labels, faces, text or moderation at scale

Amazon Rekognition

by Amazon Web Services · Usage-based · $0.001 per image
7.3/10

Amazon Rekognition is the simplest choice if you already run on AWS and want common vision tasks without training anything. Its ready-made APIs detect objects, scenes and landmarks, analyse and compare faces, read text, spot protective equipment and flag unsafe content in images and stored video.

Prices are low and clear. Image analysis costs $0.001 per image for the first million images a month, and new accounts get 1,000 free images a month in each API group for 12 months. Stored-video label detection is $0.10 per minute. For your own objects, Custom Labels trains a model from labeled images for $1 per training hour, then charges $4 per hour while the model is running. Leave one model on around the clock and that is about $2,920 a month.

AWS is trimming the product. On 30 April 2026, the Streaming Events and Batch Image Content Moderation features closed to new customers, though existing users can keep them.

Pick it if your images are already in S3 and you need faces, text or moderation at volume. Skip it if you need edge deployment or open-vocabulary detection; Rekognition runs only in AWS's cloud.

Score breakdown

Accuracy & capabilities7.5
Custom training7.0
Deployment & edge6.0
Pricing & value8.0
Developer experience8.0

Key facts

Pricing
$0.001 per image (Image APIs: $0.001 per image for the first 1 million images a month. Free tier: 1,000 images/month in each API group for 12 months. Stored video label detection: $0.10/minute. Custom Labels: $1 per training hour and $4 per inference hour.)
Free option
Yes
Platforms
API, AWS SDKs, AWS Console
Image price
$0.001 per image (first 1M/month)
Custom Labels
$1/training hour, $4/inference hour
Features
Labels, faces, text, unsafe content, PPE, celebrities
Maintenance mode
Streaming Events and Batch Image Content Moderation closed to new customers on 30 Apr 2026

What we like

  • $0.001 per image with a 12-month free tier
  • Broad ready-made features, including face search and PPE
  • Custom Labels for your own objects
  • Fits neatly into S3 and Lambda pipelines

Watch out for

  • Cloud only
  • Custom Labels endpoints bill every hour they run
  • Some features closed to new customers in 2026
#8 · Reliable OCR and image labels on Google Cloud

Google Cloud Vision API

by Google Cloud · Usage-based · Free (1,000 units/month); then $1.50 per 1,000
7.0/10

Google Cloud Vision is Google's long-running ready-made image API. With one call it can add labels, locate objects, read text (including dense documents), detect faces, logos and landmarks, and flag unsafe content with SafeSearch.

Pricing is simple. The first 1,000 units of each feature every month are free. After that, most features cost $1.50 per 1,000 images, object localization costs $2.25 and web detection $3.50, with discounts above 5 million a month. Each feature you ask for on an image is billed separately, so request only what you need.

The bigger picture is that Google is steering new vision work toward Gemini. At Cloud Next in April 2026, Vertex AI became the Gemini Enterprise Agent Platform. Vertex AI Vision, Google's managed video-analytics service, was deprecated on 15 June 2026 and ends on 30 September 2026, with Cloud Vision API named as one migration path. AutoML image training for custom models still exists inside Agent Platform as a separate product.

Pick it if you want dependable OCR or labels at predictable prices on Google Cloud. Skip it if you need your own object classes; the Gemini API or Roboflow is more flexible.

Score breakdown

Accuracy & capabilities7.5
Custom training6.5
Deployment & edge5.5
Pricing & value7.5
Developer experience8.0

Key facts

Pricing
Free (1,000 units/month); then $1.50 per 1,000 (First 1,000 units per feature per month free. Most features $1.50 per 1,000 images up to 5 million; object localization $2.25; web detection $3.50. Each feature requested on an image is billed separately.)
Free option
Yes
Platforms
API, Client libraries, Google Cloud console
Features
Labels, OCR, faces, logos, landmarks, objects, SafeSearch, web detection
Free tier
1,000 units per feature per month
Platform change
Vertex AI became Gemini Enterprise Agent Platform (April 2026)
Vertex AI Vision
Deprecated 15 Jun 2026; end of life 30 Sep 2026

What we like

  • Simple per-image pricing with a monthly free tier
  • Strong OCR, including dense document text
  • Mature client libraries and docs

Watch out for

  • Fixed label set; custom objects need a separate product
  • Cloud only
  • Every feature per image is billed separately
#9 · Existing customers running vision and language models on one platform

Clarifai

by Clarifai · Usage-based · Usage-based (current rates not confirmed)
6.2/10

Clarifai built its name on image recognition and still describes a full platform for vision work: custom classification and detection, OCR, video analysis and content moderation, plus a catalogue of vision and language models you can chain together in workflows. Its Compute Orchestration product runs models on Clarifai's GPUs, your own cloud or on-premises servers.

The reason it ranks this low is uncertainty. In May 2026 Nebius hired Clarifai's core engineering team, including founder and CEO Matthew Zeiler, took its patents and licensed its inference and orchestration technology. The report of the deal did not say what happens to the Clarifai platform or its customers. When we checked on 25 September 2026, Clarifai's pricing and documentation sites did not load for us.

On pricing, Clarifai has said it retired its old self-serve plans in favour of one Pay-As-You-Go plan billed on tokens and GPU time, with a default $100 monthly spending cap. We could not confirm current rates.

Pick it if you already run production workloads on Clarifai and have a support contract that covers its future. Skip it if you are starting fresh; Roboflow or a cloud API is the safer bet today.

Score breakdown

Accuracy & capabilities7.0
Custom training7.0
Deployment & edge6.5
Pricing & value5.0
Developer experience5.0

Key facts

Pricing
Usage-based (current rates not confirmed) (Clarifai says it replaced its old self-serve plans with one Pay-As-You-Go plan billed on tokens and GPU time, with a default $100 monthly spending cap. Its pricing page did not load when we checked on 25 September 2026, so we could not confirm current rates.)
Free option
No
Platforms
Web, API, Python SDK
Capabilities
Classification, detection, OCR, moderation, custom training
Ownership news
Nebius hired core team and CEO, licensed orchestration tech (May 2026)
Billing
Pay-as-you-go, $100 default monthly cap (Clarifai)

What we like

  • Covers vision, OCR and moderation in one platform
  • Runs models across clouds and on-premises
  • Workflow editor for chaining models

Watch out for

  • Core team and CEO moved to Nebius in May 2026
  • Current prices could not be confirmed
  • Future of the platform is unclear
#10 · Existing Azure users who need time to migrate

Azure AI Vision

by Microsoft · Usage-based · Free (5,000 transactions/month); then pay per 1,000
5.9/10

Azure AI Vision, now branded Azure Vision in Foundry Tools, is Microsoft's ready-made image API. Version 4.0 reads text, writes captions and dense captions, adds tags, detects objects and people, and suggests smart crops. Version 3.2 adds brands, faces, landmarks and adult-content checks. It also runs in containers, including disconnected ones, which some regulated customers need.

We rank it last because Microsoft is winding it down. The Image Analysis API (versions 3.2 and 4.0) will be retired on 25 September 2028, after which calls fail. Custom Vision, Microsoft's tool for training your own classifier or detector, retires the same day. The custom-model and background-removal features inside Image Analysis 4.0 were already switched off on 31 March 2025. Microsoft asked customers to have a migration plan in place by 25 September 2026.

Microsoft's suggested replacements are Document Intelligence for OCR, the Face API for faces, GPT models in Microsoft Foundry, and Azure Content Understanding for managed image analysis.

Pick it if you already run it in production and need a stable bridge while you migrate. Skip it for any new project; build on Foundry models, Roboflow or Ultralytics instead.

Score breakdown

Accuracy & capabilities7.0
Custom training3.5
Deployment & edge6.5
Pricing & value6.5
Developer experience5.5

Key facts

Pricing
Free (5,000 transactions/month); then pay per 1,000 (Free tier: 5,000 transactions a month at 20 per minute. Paid rates per 1,000 transactions are shown in the Azure pricing calculator for your region and agreement.)
Free option
Yes
Platforms
API, SDKs, Containers (connected and disconnected)
Now called
Azure Vision in Foundry Tools
Image Analysis retires
25 September 2028 (v3.2 and v4.0)
Custom Vision retires
25 September 2028
Already retired
Image Analysis 4.0 custom models and background removal (31 Mar 2025)

What we like

  • Free tier of 5,000 transactions a month
  • Runs in disconnected containers
  • Captions, OCR and people detection in one call

Watch out for

  • Image Analysis and Custom Vision retire on 25 Sep 2028
  • Custom-model features already removed
  • Paid prices not shown on the public pricing page

How we scored these tools

Each tool is scored 0–10 on the criteria below, using public evidence: independent benchmarks, vendor documentation and pricing pages, aggregate user ratings and reputable reviews. The overall score is the weighted average. Nobody pays to be listed. Read our full methodology.

CriterionWeightWhat we look at
Accuracy & capabilities25%Published benchmark results and the range of tasks covered: classification, detection, segmentation, OCR, pose and video tracking.
Custom training20%How easily you can teach it your own objects, from labeled images or prompts, and how much data that needs.
Deployment & edge20%Cloud API, self-hosting, on-device and offline options, export formats, and whether the product has a stable future.
Pricing & value20%Free tier, published prices, licence costs for commercial use and how bills grow at scale.
Developer experience15%SDKs, docs, time to a first working result and how much glue code you must write.

Four ways to build a vision app in 2026

Before you pick a tool, pick an approach. Most projects use two of these together.

Approach Examples Good for Watch out for
Train your own model Roboflow, Ultralytics YOLO, LandingLens Specific objects, real-time video, edge devices You need labeled images (hundreds to thousands)
Ready-made cloud API Amazon Rekognition, Google Cloud Vision, Azure AI Vision Common tasks: OCR, labels, faces, moderation Fixed label sets, cloud only, some products retiring
Open foundation model Meta SAM 3 Outlining and tracking any object from a text prompt Needs a strong GPU; custom licence
Multimodal LLM Gemini API Prototypes, many object types, low volume Slower, less consistent boxes, no offline use

A common 2026 pattern: prompt Gemini or SAM 3 to pre-label a few thousand images, have people correct the labels, then train a small YOLO or RF-DETR model that runs cheaply on your own hardware. OpenCV often handles the video input and output around that model. For labeling tools, see our best data labeling tools ranking.

Pricing guide (as of 25 September 2026)

Tool Free option Paid entry What you pay for
Roboflow Public plan (data is public) $79/month Core, billed yearly Credits for labeling, training, inference
Ultralytics YOLO AGPL-3.0 code; Platform $25 one-time credit $29/seat/month Pro Seats, GPU hours from $0.24; Enterprise License for closed products
Meta SAM 3 Free weights None Your own GPUs
OpenCV Free (Apache 2.0) None Nothing
Gemini API Free tier $0.75 per 1M input tokens (3.8 Flash) Tokens (258 per small image)
LandingLens 1,000 credits/month, non-commercial Enterprise (custom) Credits per image
Amazon Rekognition 1,000 images/month for 12 months $0.001 per image Images, video minutes, Custom Labels hours
Google Cloud Vision 1,000 units/month per feature $1.50 per 1,000 Each feature per image
Clarifai Not confirmed Pay-as-you-go Tokens and GPU time
Azure AI Vision 5,000 transactions/month Per 1,000 transactions Transactions

Worked example: 100,000 images a month through Rekognition label detection costs 100,000 x $0.001 = $100 (after any free tier). The same volume through Google Cloud Vision label detection costs 99 x $1.50 = about $149, because the first 1,000 are free. A self-hosted YOLO or RF-DETR model on a GPU you already own costs only electricity.

Licences: what you are allowed to ship

The licence matters more than benchmark points if you sell a product.

  • Apache 2.0 (OpenCV, Roboflow Inference core, RF-DETR core models): use it commercially, keep your own code private, just keep the notices.
  • AGPL-3.0 (Ultralytics YOLO by default): if people use your product, including over a network, you must share your source code under the same licence. Most companies either comply or buy Ultralytics' Enterprise License.
  • Custom model licences (Meta's SAM License): commercial use is allowed with some restrictions. Read the terms before you ship.
  • Cloud APIs: no licence to manage, but your images go to the vendor and you depend on the product staying alive.

If you are unsure, ask your legal team before training on top of an AGPL model, because models you train inherit that licence by default.

The big clouds are retreating from classic vision APIs

2026 is a year of cleanup for older vision services.

  • Microsoft: the Image Analysis API and Custom Vision both retire on 25 September 2028. Microsoft points customers to Document Intelligence, the Face API, GPT models in Foundry and Azure Content Understanding.
  • Google: Vertex AI became the Gemini Enterprise Agent Platform in April 2026. Vertex AI Vision ends on 30 September 2026. Cloud Vision API continues.
  • AWS: Rekognition's Streaming Events and Batch Image Content Moderation features closed to new customers on 30 April 2026.

The direction is clear: the clouds want new vision work to run on general multimodal models. That is fine for many tasks, but real-time and offline use still needs a small trained model. Build with portable formats such as ONNX so you can move if a service is retired.

How we ranked these tools

We scored each tool from 0 to 10 on five criteria: accuracy and capabilities (25%), custom training (20%), deployment and edge (20%), pricing and value (20%) and developer experience (15%). The overall score is the weighted average. A product's announced retirement lowers its deployment score, because you cannot build on a service that is going away.

We used public sources only: vendor pricing pages and docs, model cards and GitHub repositories, official deprecation notices and reputable press. Benchmark numbers such as COCO mAP and speed claims are the vendors' own and are labelled as such. We did not run our own tests and did not accept payment for placement.

Expert tips
  1. Before you label anything, send 50 of your real images to the Gemini API or SAM 3 with a plain prompt. If the results are good enough, you may not need to train a model at all. If they are not, keep the outputs as pre-labels to fix by hand.
  2. Test on your own camera footage, not benchmark scores. Hold back images from a different day, site or lighting condition, because models that score well on COCO often slip when the camera angle changes.
  3. Check the licence before your first training run. Models trained with Ultralytics YOLO are AGPL-3.0 by default; if you cannot open-source your product, budget for an Enterprise License or use Apache 2.0 RF-DETR.
  4. Stop cloud endpoints you are not using. A Rekognition Custom Labels model costs $4 an hour while it runs, about $2,920 a month if left on.
  5. Export trained models to ONNX. It keeps you portable between OpenCV 5, TensorRT, OpenVINO and cloud services if a vendor changes prices or retires a product.

Jargon explained

Object detection
Finding objects in an image and drawing a labeled box around each one, for example every car in a parking lot photo.
Segmentation
Marking the exact pixels that belong to each object, instead of a rough box. Useful when shape and size matter.
mAP (mean average precision)
The standard accuracy score for detectors on benchmarks such as COCO. Higher is better; it rewards finding objects and placing boxes accurately.
Edge deployment
Running a model on a device near the camera, such as a phone, a Raspberry Pi or an NVIDIA Jetson, instead of sending images to the cloud.
Open-vocabulary
A model that can find objects named in a text prompt, even if it was never trained on that exact label list.
AGPL-3.0
An open-source licence that requires you to share your own source code if you distribute the software or let people use it over a network.

Frequently asked questions

What is the best computer vision tool in 2026?

For most teams, Roboflow. It handles labeling, training and deployment in one product, runs on edge devices, and private projects start at $79/month billed yearly. Developers who prefer code should look at Ultralytics YOLO, and anyone who wants no training at all can start with the Gemini API.

Is YOLO free for commercial use?

Ultralytics YOLO is free under the AGPL-3.0 licence, which requires you to publish your own source code if you ship it in a product. To keep your code private you need Ultralytics' paid Enterprise License, which has no public price. Roboflow's RF-DETR core models are Apache 2.0 and have no such requirement.

Can Gemini or ChatGPT replace a trained object detector?

For prototypes and low volumes, often yes. Gemini can return bounding boxes and segmentation masks from a plain-English prompt. For real-time video, offline use or consistent pixel-level accuracy, a small model trained on your own images (YOLO or RF-DETR) is still faster, cheaper per image and more predictable.

What is the best free computer vision tool?

OpenCV for image and video processing (Apache 2.0), Meta SAM 3 for segmentation from a text prompt, and RF-DETR or Ultralytics YOLO for object detection you train yourself. YOLO is free only if you accept the AGPL-3.0 terms.

What is replacing Azure Custom Vision?

Microsoft will retire Custom Vision on 25 September 2028. It recommends Azure Machine Learning AutoML for training custom classifiers and detectors, or generative models in Microsoft Foundry and Azure Content Understanding. Outside Azure, Roboflow and Ultralytics are the most direct replacements.

How much does a computer vision API cost?

Ready-made APIs cost roughly $1 to $1.50 per 1,000 images: Amazon Rekognition is $0.001 per image, and most Google Cloud Vision features are $1.50 per 1,000 after 1,000 free each month. Always-on custom models cost more; a Rekognition Custom Labels model bills $4 for every hour it runs.

What is the difference between object detection and segmentation?

Object detection draws a box around each object and names it. Segmentation outlines the exact pixels that belong to each object, which is more precise and useful for measuring size, removing backgrounds or guiding robots. SAM 3 specialises in segmentation; YOLO and RF-DETR do both.

Sources

Every fact on this page comes from public information. Vendor figures are labelled as vendor claims.

  1. Roboflow pricing (Roboflow)
  2. RF-DETR GitHub repository (GitHub)
  3. RF-DETR: a SOTA real-time object detection model (Roboflow)
  4. Roboflow Inference GitHub repository (GitHub)
  5. Ultralytics pricing (Ultralytics)
  6. Ultralytics YOLO26 docs (Ultralytics)
  7. Ultralytics licensing (Ultralytics)
  8. SAM 3.1: faster real-time video detection and tracking (Meta AI)
  9. Meta releases SAM 3 (InfoQ)
  10. SAM 3 GitHub repository (GitHub)
  11. OpenCV 5.0 released with rewritten DNN engine (Phoronix)
  12. OpenCV 5 release: new DNN engine with enhanced ONNX and LLM/VLM support (CNX Software)
  13. OpenCV releases (GitHub)
  14. Gemini API image understanding (Google)
  15. Gemini API pricing (Google)
  16. LandingLens plans (LandingAI)
  17. LandingAI pricing (LandingAI)
  18. Amazon Rekognition pricing (AWS)
  19. Amazon Rekognition image features (AWS)
  20. AWS service availability updates (March 2026) (AWS)
  21. Cloud Vision API pricing (Google Cloud)
  22. Vertex AI Vision deprecation notice (Google Cloud)
  23. Gemini Enterprise Agent Platform (formerly Vertex AI) (Google Cloud)
  24. Nebius snaps up Clarifai's compute orchestration tech and talent (SiliconANGLE)
  25. Clarifai: introducing pay-as-you-go credits (Clarifai)
  26. What is Image Analysis? (Azure Vision in Foundry Tools) (Microsoft Learn)
  27. Migrate from Azure Vision Image Analysis (Microsoft Learn)
  28. What's new in Custom Vision (retirement notice) (Microsoft Learn)
  29. Azure AI Vision pricing (Microsoft)