Developers & AI InfrastructureBest Computer Vision Tools
Also known as: computer vision platforms
Computer vision tools let software understand images and video: find objects, read text, count products on a shelf or spot a scratch on a production line. This list covers three kinds. Platforms (Roboflow, Ultralytics, LandingLens) help you label images, train your own model and deploy it. Ready-made cloud APIs (Amazon Rekognition, Google Cloud Vision, Azure AI Vision) recognise common things like labels, faces and text with no training. Open models and libraries (OpenCV, Meta's SAM 3) are free to run on your own hardware.
A fourth option now competes with all three: multimodal AI models such as Gemini can find and outline objects from a plain-English prompt. We ranked 10 tools on accuracy, custom training, deployment, price and developer experience. Prices are as of 25 September 2026. Note that two big clouds are pulling back: Microsoft will retire Azure's Image Analysis and Custom Vision APIs on 25 September 2028, and Google's Vertex AI Vision reaches end of life on 30 September 2026.
Quick answerRoboflow is the best computer vision tool for most teams in 2026. It covers labeling, training, deployment and edge inference in one place, has a free plan, and private projects start at $79/month billed yearly. Pick Ultralytics YOLO to train small, fast detectors yourself (check the AGPL-3.0 licence first), SAM 3 for free open-vocabulary segmentation, and the Gemini API when a prompt like "find every dented can" is good enough and you do not want to train anything.
Top picks at a glance
Best overallRoboflow
Labeling, training and deployment in one platform, with an Apache 2.0 detector (RF-DETR) and inference server you can self-host.
Best for edge-ready detectorsUltralytics YOLO
YOLO26 trains in a few lines of Python and exports to ONNX, TensorRT, CoreML, LiteRT and OpenVINO.
Best for segmentationMeta SAM 3
Free model that outlines and tracks every object matching a text prompt, in images and video.
Best with no trainingGemini API (vision)
Returns boxes and outlines from a plain-English request, with a free tier and low per-image cost.
Best ready-made cloud APIAmazon Rekognition
Labels, faces, text and moderation at $0.001 per image, plus Custom Labels for your own objects.
Scoreboard
Scores are out of 10. The overall score is the weighted average of the criteria below.
| # | Tool | Overall | Accuracy & capabilities | Custom training | Deployment & edge | Pricing & value | Developer experience | Price from | Best for |
|---|
| 1 | Roboflow Roboflow | 8.9 | 8.8 | 9.5 | 9.0 | 8.0 | 9.3 | Free (public data); $79/month Core (billed yearly) Free tier | Teams that need a custom detector or segmenter in production quickly |
| 2 | Ultralytics YOLO Ultralytics | 8.8 | 8.8 | 9.0 | 9.5 | 8.0 | 8.8 | Free (AGPL-3.0); Platform Pro $29/seat/month Free tier | Python developers training fast detectors for cameras and edge devices |
| 3 | Meta SAM 3 Meta | 7.9 | 9.0 | 6.5 | 7.0 | 9.0 | 7.5 | Free Free tier | Precise outlines, auto-labeling and video tracking from a text prompt |
| 4 | OpenCV OpenCV (open-source project) | 7.7 | 7.0 | 5.0 | 9.5 | 10.0 | 7.0 | Free Free tier | Engineers building image and video pipelines with zero licence cost |
| 5 | Gemini API (vision) Google | 7.6 | 8.5 | 6.0 | 6.0 | 8.5 | 9.0 | Free tier; then $0.75 per 1M input tokens (Gemini 3.8 Flash) Free tier | Prototypes and low-volume jobs where a prompt can replace a trained model |
| 6 | LandingLens LandingAI | 7.5 | 7.5 | 8.5 | 8.0 | 5.5 | 8.0 | Free (non-commercial); Enterprise custom Free tier | Manufacturers running visual quality inspection |
| 7 | Amazon Rekognition Amazon Web Services | 7.3 | 7.5 | 7.0 | 6.0 | 8.0 | 8.0 | $0.001 per image Free tier | AWS teams needing labels, faces, text or moderation at scale |
| 8 | Google Cloud Vision API Google Cloud | 7.0 | 7.5 | 6.5 | 5.5 | 7.5 | 8.0 | Free (1,000 units/month); then $1.50 per 1,000 Free tier | Reliable OCR and image labels on Google Cloud |
| 9 | Clarifai Clarifai | 6.2 | 7.0 | 7.0 | 6.5 | 5.0 | 5.0 | Usage-based (current rates not confirmed) | Existing customers running vision and language models on one platform |
| 10 | Azure AI Vision Microsoft | 5.9 | 7.0 | 3.5 | 6.5 | 6.5 | 5.5 | Free (5,000 transactions/month); then pay per 1,000 Free tier | Existing Azure users who need time to migrate |
Expert reviews
Roboflow is the quickest way to go from a folder of photos to a working vision model. It covers the whole loop: upload and label images (with Segment Anything-based tools and auto-labeling), train a model on hosted GPUs, test it, then deploy it to a hosted API or your own hardware.
Its own model family, RF-DETR, is open source. Roboflow says it was the first real-time detector to pass 60 mAP on the COCO benchmark, the research paper was accepted at ICLR 2026, and segmentation and keypoint versions followed this year. The Inference server that runs models is also Apache 2.0 and works on NVIDIA Jetson, Raspberry Pi, CPUs and GPUs, so you are not locked into Roboflow's cloud.
Pricing is published. The free Public plan gives 15 credits a month but makes your data and models public. Core costs $79/month billed yearly ($99 monthly) for three users and private data. Credits pay for labeling, training and inference, so heavy users should budget for add-ons.
Pick it if you need a custom detector in days, not months. Skip it if your images must stay private and you have no budget; train locally with Ultralytics or RF-DETR instead.
What we like
- Labeling, training and deployment in one product
- Apache 2.0 RF-DETR models and self-hostable Inference server
- Runs on edge devices such as Jetson and Raspberry Pi
- Clear published prices
Watch out for
- Free plan makes your data and models public
- Credits cap how much you can train and run on Core
- Some RF-DETR Plus components use a non-Apache licence
Ultralytics makes YOLO, one of the best-known families of real-time object detectors, and its newest version, YOLO26, arrived in January 2026. One Python package covers detection, instance and semantic segmentation, classification, pose, rotated boxes and depth estimation. Training on your own labeled images takes a few lines of code.
The numbers are strong for the size. Ultralytics reports 40.9 mAP on COCO for the tiny YOLO26n and 57.5 for YOLO26x, and says YOLO26n runs up to 43% faster on CPU than YOLO11n. Models export to ONNX, TensorRT, CoreML, LiteRT and OpenVINO, so they run on phones, Jetson boards and plain CPUs.
The catch is the licence. The code and the models you train are AGPL-3.0 by default. If you ship YOLO inside a commercial product without open-sourcing your own code, you need an Enterprise License, and its price is not public. The hosted Ultralytics Platform starts free with a $25 one-time credit; Pro is $29 per seat per month plus GPU time.
Pick it if you can write Python and want the shortest path to an edge-ready detector. Skip it if AGPL is a problem and you will not buy a licence; RF-DETR's core models are Apache 2.0.
What we like
- Seven vision tasks from one package and one API
- Small models that run fast on CPUs and edge devices
- Exports to every major edge runtime
- Cheap hosted training from $0.24/GPU-hour
Watch out for
- AGPL-3.0 by default; closed commercial use needs a paid licence
- Enterprise License price is not published
- You still need labeled data from somewhere
SAM 3 is Meta's third Segment Anything model, released in November 2025, and it changed what a free model can do. Earlier versions outlined whatever you clicked on. SAM 3 takes a short text prompt such as "yellow school bus", or an example image, and then finds, outlines and tracks every matching object in a photo or video. This is called open-vocabulary segmentation: there is no fixed list of classes and no training step.
Meta says SAM 3 doubled the accuracy of earlier systems on its SA-Co benchmark and takes about 30 milliseconds per image on an H200 GPU, even with more than 100 objects. SAM 3.1 (March 2026) is a drop-in update that tracks up to 16 objects in one pass and, according to Meta, doubles video throughput to 32 frames per second on one H100.
It is free to download under Meta's own SAM License, which allows commercial use with some restrictions, so read it before you ship. You need a capable GPU, and it gives you masks, not a finished app.
Pick it if you need precise outlines, fast auto-labeling or video tracking. Skip it if you need a tiny model for a phone; train a YOLO or RF-DETR model instead, using SAM 3 to label its data.
What we like
- Finds and outlines every object matching a text prompt
- Tracks objects through video
- Free weights, with commercial use allowed
- Excellent for pre-labeling training data
Watch out for
- Needs a data-center or strong desktop GPU for good speed
- Custom licence, not a standard open-source one
- Outputs masks only; you build the rest
OpenCV is the free, open-source library that many vision projects rely on for the plumbing: reading cameras and video files, resizing and colour conversion, filters, feature matching, camera calibration, 3D geometry and drawing results on screen. It works from C++, Python, Java and JavaScript, on hardware from servers to single-board computers.
OpenCV 5.0 shipped in June 2026 with a rewritten deep-learning (DNN) engine. Support for ONNX operators rose from about 22% in version 4 to over 80%, so many more modern models load directly, and the release adds building blocks for running language and vision-language models. It also dropped the old C API and now needs C++17, so older code may need porting. The licence is Apache 2.0.
What OpenCV does not do is train modern neural networks, label data or give you a dashboard. You bring a trained model (from Ultralytics, Roboflow or PyTorch), and OpenCV runs it and handles everything around it.
Pick it if you are an engineer building a pipeline and want no licence fees or vendor lock-in. Skip it if you want a no-code tool or a hosted API; choose Roboflow or a cloud API instead.
What we like
- Free under Apache 2.0, with no usage limits
- Runs almost anywhere, including phones and small boards
- New DNN engine loads far more ONNX models
- Huge community and decades of tutorials
Watch out for
- No model training, labeling or hosting
- Version 5 breaks some old C and C++ code
- Needs programming skills
Multimodal AI models now handle jobs that once needed a custom-trained detector, and Google's Gemini API is the one we would try first, because its docs support object detection with bounding boxes and segmentation outlines out of the box. You send an image and a plain-English request, such as "find every dented can", and it returns labels, boxes or masks as structured data. It also reads text, answers questions about photos and understands video.
Google's docs recommend Gemini 3.8 Flash for image understanding. It costs $0.75 per million input tokens and $3.75 per million output tokens until 31 December 2026, then doubles, and there is a free tier. A small image counts as 258 tokens, so the input for one image costs a tiny fraction of a cent; the answer adds more.
The trade-offs are real. A general model is not tuned to your objects, so expect less consistent boxes than a detector trained on your own images. Each call also travels to Google's cloud, so it will not work offline or keep up with a 30-frames-per-second camera.
Pick it if you need a prototype today, handle many object types or process low volumes. Skip it if you need real-time, offline or pixel-exact results. See our Gemini 3.8 Flash page.
What we like
- No training: describe what to find in plain English
- Boxes and segmentation masks returned as JSON
- Free tier and low per-image cost
- Handles images, documents and video in one API
Watch out for
- Cloud only, with no offline or on-device option
- Too slow for real-time camera streams
- Less consistent than a detector trained on your data
- Price doubles on 1 January 2027
LandingLens is LandingAI's platform for visual inspection in factories: finding scratches, dents, missing parts and other defects on a production line. You label images, train a model and deploy it from one guided web app, so quality engineers can own the project without a machine-learning team.
Deployment fits factory floors well. Models can run in LandingAI's cloud, on site through the LandingEdge app, or in a Docker container. The free plan is usable for trials: 1,000 credits a month, up to 3 invited users, unlimited projects and 10,000 images per project. With Fast Training, training on or running one image costs 1 credit.
The limits are commercial. The free plan's single model download is for non-commercial use only, and production use needs the Enterprise plan, which has no public price. LandingAI's main pricing page now covers only its document product, Agentic Document Extraction, which suggests where the company's attention has moved.
Pick it if you run quality inspection and want a guided tool your engineers can manage. Skip it if you need general-purpose detection or published prices; Roboflow is more open about cost.
What we like
- Built around factory defect inspection
- Edge and Docker deployment for on-site lines
- Guided workflow for non-specialists
- Free plan for trials
Watch out for
- Commercial use requires a custom-priced Enterprise plan
- Narrower than general vision platforms
- Vendor's focus has shifted toward document extraction
Amazon Rekognition is the simplest choice if you already run on AWS and want common vision tasks without training anything. Its ready-made APIs detect objects, scenes and landmarks, analyse and compare faces, read text, spot protective equipment and flag unsafe content in images and stored video.
Prices are low and clear. Image analysis costs $0.001 per image for the first million images a month, and new accounts get 1,000 free images a month in each API group for 12 months. Stored-video label detection is $0.10 per minute. For your own objects, Custom Labels trains a model from labeled images for $1 per training hour, then charges $4 per hour while the model is running. Leave one model on around the clock and that is about $2,920 a month.
AWS is trimming the product. On 30 April 2026, the Streaming Events and Batch Image Content Moderation features closed to new customers, though existing users can keep them.
Pick it if your images are already in S3 and you need faces, text or moderation at volume. Skip it if you need edge deployment or open-vocabulary detection; Rekognition runs only in AWS's cloud.
What we like
- $0.001 per image with a 12-month free tier
- Broad ready-made features, including face search and PPE
- Custom Labels for your own objects
- Fits neatly into S3 and Lambda pipelines
Watch out for
- Cloud only
- Custom Labels endpoints bill every hour they run
- Some features closed to new customers in 2026
Google Cloud Vision is Google's long-running ready-made image API. With one call it can add labels, locate objects, read text (including dense documents), detect faces, logos and landmarks, and flag unsafe content with SafeSearch.
Pricing is simple. The first 1,000 units of each feature every month are free. After that, most features cost $1.50 per 1,000 images, object localization costs $2.25 and web detection $3.50, with discounts above 5 million a month. Each feature you ask for on an image is billed separately, so request only what you need.
The bigger picture is that Google is steering new vision work toward Gemini. At Cloud Next in April 2026, Vertex AI became the Gemini Enterprise Agent Platform. Vertex AI Vision, Google's managed video-analytics service, was deprecated on 15 June 2026 and ends on 30 September 2026, with Cloud Vision API named as one migration path. AutoML image training for custom models still exists inside Agent Platform as a separate product.
Pick it if you want dependable OCR or labels at predictable prices on Google Cloud. Skip it if you need your own object classes; the Gemini API or Roboflow is more flexible.
What we like
- Simple per-image pricing with a monthly free tier
- Strong OCR, including dense document text
- Mature client libraries and docs
Watch out for
- Fixed label set; custom objects need a separate product
- Cloud only
- Every feature per image is billed separately
Clarifai built its name on image recognition and still describes a full platform for vision work: custom classification and detection, OCR, video analysis and content moderation, plus a catalogue of vision and language models you can chain together in workflows. Its Compute Orchestration product runs models on Clarifai's GPUs, your own cloud or on-premises servers.
The reason it ranks this low is uncertainty. In May 2026 Nebius hired Clarifai's core engineering team, including founder and CEO Matthew Zeiler, took its patents and licensed its inference and orchestration technology. The report of the deal did not say what happens to the Clarifai platform or its customers. When we checked on 25 September 2026, Clarifai's pricing and documentation sites did not load for us.
On pricing, Clarifai has said it retired its old self-serve plans in favour of one Pay-As-You-Go plan billed on tokens and GPU time, with a default $100 monthly spending cap. We could not confirm current rates.
Pick it if you already run production workloads on Clarifai and have a support contract that covers its future. Skip it if you are starting fresh; Roboflow or a cloud API is the safer bet today.
What we like
- Covers vision, OCR and moderation in one platform
- Runs models across clouds and on-premises
- Workflow editor for chaining models
Watch out for
- Core team and CEO moved to Nebius in May 2026
- Current prices could not be confirmed
- Future of the platform is unclear
Azure AI Vision, now branded Azure Vision in Foundry Tools, is Microsoft's ready-made image API. Version 4.0 reads text, writes captions and dense captions, adds tags, detects objects and people, and suggests smart crops. Version 3.2 adds brands, faces, landmarks and adult-content checks. It also runs in containers, including disconnected ones, which some regulated customers need.
We rank it last because Microsoft is winding it down. The Image Analysis API (versions 3.2 and 4.0) will be retired on 25 September 2028, after which calls fail. Custom Vision, Microsoft's tool for training your own classifier or detector, retires the same day. The custom-model and background-removal features inside Image Analysis 4.0 were already switched off on 31 March 2025. Microsoft asked customers to have a migration plan in place by 25 September 2026.
Microsoft's suggested replacements are Document Intelligence for OCR, the Face API for faces, GPT models in Microsoft Foundry, and Azure Content Understanding for managed image analysis.
Pick it if you already run it in production and need a stable bridge while you migrate. Skip it for any new project; build on Foundry models, Roboflow or Ultralytics instead.
What we like
- Free tier of 5,000 transactions a month
- Runs in disconnected containers
- Captions, OCR and people detection in one call
Watch out for
- Image Analysis and Custom Vision retire on 25 Sep 2028
- Custom-model features already removed
- Paid prices not shown on the public pricing page
How we scored these tools
Each tool is scored 0–10 on the criteria below, using public evidence: independent benchmarks, vendor documentation and pricing pages, aggregate user ratings and reputable reviews. The overall score is the weighted average. Nobody pays to be listed. Read our full methodology.
| Criterion | Weight | What we look at |
|---|
| Accuracy & capabilities | 25% | Published benchmark results and the range of tasks covered: classification, detection, segmentation, OCR, pose and video tracking. |
| Custom training | 20% | How easily you can teach it your own objects, from labeled images or prompts, and how much data that needs. |
| Deployment & edge | 20% | Cloud API, self-hosting, on-device and offline options, export formats, and whether the product has a stable future. |
| Pricing & value | 20% | Free tier, published prices, licence costs for commercial use and how bills grow at scale. |
| Developer experience | 15% | SDKs, docs, time to a first working result and how much glue code you must write. |
Four ways to build a vision app in 2026
Before you pick a tool, pick an approach. Most projects use two of these together.
| Approach |
Examples |
Good for |
Watch out for |
| Train your own model |
Roboflow, Ultralytics YOLO, LandingLens |
Specific objects, real-time video, edge devices |
You need labeled images (hundreds to thousands) |
| Ready-made cloud API |
Amazon Rekognition, Google Cloud Vision, Azure AI Vision |
Common tasks: OCR, labels, faces, moderation |
Fixed label sets, cloud only, some products retiring |
| Open foundation model |
Meta SAM 3 |
Outlining and tracking any object from a text prompt |
Needs a strong GPU; custom licence |
| Multimodal LLM |
Gemini API |
Prototypes, many object types, low volume |
Slower, less consistent boxes, no offline use |
A common 2026 pattern: prompt Gemini or SAM 3 to pre-label a few thousand images, have people correct the labels, then train a small YOLO or RF-DETR model that runs cheaply on your own hardware. OpenCV often handles the video input and output around that model. For labeling tools, see our best data labeling tools ranking.
Pricing guide (as of 25 September 2026)
| Tool |
Free option |
Paid entry |
What you pay for |
| Roboflow |
Public plan (data is public) |
$79/month Core, billed yearly |
Credits for labeling, training, inference |
| Ultralytics YOLO |
AGPL-3.0 code; Platform $25 one-time credit |
$29/seat/month Pro |
Seats, GPU hours from $0.24; Enterprise License for closed products |
| Meta SAM 3 |
Free weights |
None |
Your own GPUs |
| OpenCV |
Free (Apache 2.0) |
None |
Nothing |
| Gemini API |
Free tier |
$0.75 per 1M input tokens (3.8 Flash) |
Tokens (258 per small image) |
| LandingLens |
1,000 credits/month, non-commercial |
Enterprise (custom) |
Credits per image |
| Amazon Rekognition |
1,000 images/month for 12 months |
$0.001 per image |
Images, video minutes, Custom Labels hours |
| Google Cloud Vision |
1,000 units/month per feature |
$1.50 per 1,000 |
Each feature per image |
| Clarifai |
Not confirmed |
Pay-as-you-go |
Tokens and GPU time |
| Azure AI Vision |
5,000 transactions/month |
Per 1,000 transactions |
Transactions |
Worked example: 100,000 images a month through Rekognition label detection costs 100,000 x $0.001 = $100 (after any free tier). The same volume through Google Cloud Vision label detection costs 99 x $1.50 = about $149, because the first 1,000 are free. A self-hosted YOLO or RF-DETR model on a GPU you already own costs only electricity.
Licences: what you are allowed to ship
The licence matters more than benchmark points if you sell a product.
- Apache 2.0 (OpenCV, Roboflow Inference core, RF-DETR core models): use it commercially, keep your own code private, just keep the notices.
- AGPL-3.0 (Ultralytics YOLO by default): if people use your product, including over a network, you must share your source code under the same licence. Most companies either comply or buy Ultralytics' Enterprise License.
- Custom model licences (Meta's SAM License): commercial use is allowed with some restrictions. Read the terms before you ship.
- Cloud APIs: no licence to manage, but your images go to the vendor and you depend on the product staying alive.
If you are unsure, ask your legal team before training on top of an AGPL model, because models you train inherit that licence by default.
The big clouds are retreating from classic vision APIs
2026 is a year of cleanup for older vision services.
- Microsoft: the Image Analysis API and Custom Vision both retire on 25 September 2028. Microsoft points customers to Document Intelligence, the Face API, GPT models in Foundry and Azure Content Understanding.
- Google: Vertex AI became the Gemini Enterprise Agent Platform in April 2026. Vertex AI Vision ends on 30 September 2026. Cloud Vision API continues.
- AWS: Rekognition's Streaming Events and Batch Image Content Moderation features closed to new customers on 30 April 2026.
The direction is clear: the clouds want new vision work to run on general multimodal models. That is fine for many tasks, but real-time and offline use still needs a small trained model. Build with portable formats such as ONNX so you can move if a service is retired.
Expert tips- Before you label anything, send 50 of your real images to the Gemini API or SAM 3 with a plain prompt. If the results are good enough, you may not need to train a model at all. If they are not, keep the outputs as pre-labels to fix by hand.
- Test on your own camera footage, not benchmark scores. Hold back images from a different day, site or lighting condition, because models that score well on COCO often slip when the camera angle changes.
- Check the licence before your first training run. Models trained with Ultralytics YOLO are AGPL-3.0 by default; if you cannot open-source your product, budget for an Enterprise License or use Apache 2.0 RF-DETR.
- Stop cloud endpoints you are not using. A Rekognition Custom Labels model costs $4 an hour while it runs, about $2,920 a month if left on.
- Export trained models to ONNX. It keeps you portable between OpenCV 5, TensorRT, OpenVINO and cloud services if a vendor changes prices or retires a product.
Jargon explained
- Object detection
- Finding objects in an image and drawing a labeled box around each one, for example every car in a parking lot photo.
- Segmentation
- Marking the exact pixels that belong to each object, instead of a rough box. Useful when shape and size matter.
- mAP (mean average precision)
- The standard accuracy score for detectors on benchmarks such as COCO. Higher is better; it rewards finding objects and placing boxes accurately.
- Edge deployment
- Running a model on a device near the camera, such as a phone, a Raspberry Pi or an NVIDIA Jetson, instead of sending images to the cloud.
- Open-vocabulary
- A model that can find objects named in a text prompt, even if it was never trained on that exact label list.
- AGPL-3.0
- An open-source licence that requires you to share your own source code if you distribute the software or let people use it over a network.
Frequently asked questions
What is the best computer vision tool in 2026?
For most teams, Roboflow. It handles labeling, training and deployment in one product, runs on edge devices, and private projects start at $79/month billed yearly. Developers who prefer code should look at Ultralytics YOLO, and anyone who wants no training at all can start with the Gemini API.
Is YOLO free for commercial use?
Ultralytics YOLO is free under the AGPL-3.0 licence, which requires you to publish your own source code if you ship it in a product. To keep your code private you need Ultralytics' paid Enterprise License, which has no public price. Roboflow's RF-DETR core models are Apache 2.0 and have no such requirement.
Can Gemini or ChatGPT replace a trained object detector?
For prototypes and low volumes, often yes. Gemini can return bounding boxes and segmentation masks from a plain-English prompt. For real-time video, offline use or consistent pixel-level accuracy, a small model trained on your own images (YOLO or RF-DETR) is still faster, cheaper per image and more predictable.
What is the best free computer vision tool?
OpenCV for image and video processing (Apache 2.0), Meta SAM 3 for segmentation from a text prompt, and RF-DETR or Ultralytics YOLO for object detection you train yourself. YOLO is free only if you accept the AGPL-3.0 terms.
What is replacing Azure Custom Vision?
Microsoft will retire Custom Vision on 25 September 2028. It recommends Azure Machine Learning AutoML for training custom classifiers and detectors, or generative models in Microsoft Foundry and Azure Content Understanding. Outside Azure, Roboflow and Ultralytics are the most direct replacements.
How much does a computer vision API cost?
Ready-made APIs cost roughly $1 to $1.50 per 1,000 images: Amazon Rekognition is $0.001 per image, and most Google Cloud Vision features are $1.50 per 1,000 after 1,000 free each month. Always-on custom models cost more; a Rekognition Custom Labels model bills $4 for every hour it runs.
What is the difference between object detection and segmentation?
Object detection draws a box around each object and names it. Segmentation outlines the exact pixels that belong to each object, which is more precise and useful for measuring size, removing backgrounds or guiding robots. SAM 3 specialises in segmentation; YOLO and RF-DETR do both.