Start / Blog / Anleitungen / YOLO-NAS Finetuning mit eigenen OMR Annotationen

YOLO-NAS Finetuning mit eigenen OMR Annotationen

Zusammenfassen mit ChatGPT

Die Objekterkennung ist eine der aufregendsten und vielseitigsten Anwendungen der künstlichen Intelligenz. Von der Bildverarbeitung über autonome Fahrzeuge bis hin zur Sicherheitsüberwachung – die Fähigkeit, Objekte in Bildern genau zu erkennen, bietet zahlreiche Möglichkeiten und Vorteile.

In diesem Artikel zeigen wir Ihnen, wie Sie mit der neuesten Technologie im Bereich der Objekterkennung, dem YOLO-NAS-Modell, ein leistungsfähiges Erkennungsmodell trainieren und einsetzen können.

Ziel des Tutorials

Diese Anleitung zeigt, wie Sie ein OMR-Modell auf annotierten Daten im COCO Format trainieren.

Wir werden das Modell für die Erkennung von Kontrollkästchen in Dokumentenstrukturen einsetzen – Sie können das Modell jedoch auch für andere Anwendungsfälle anpassen.

Voraussetzungen für die Objekterkennung

  • Allgemeines Verständnis von Objekterkennung und Training neuronaler Netzwerke
  • Ein COCO-formatiertes Objekterkennungs-Dataset
  • Umgebung zur Entwicklung von Objekterkennungs-Modellen

Um schnell loszulegen, können Sie ein neues Google Colab-Notebook erstellen oder den Code lokal oder in einer Umgebung Ihrer Wahl ausführen.

Überblick über das YOLO-NAS Modell für die Objekterkennung

Wir verwenden das hochmoderne Objekterkennungs-Modell YOLO-NAS. Dieses Modell kombiniert hohe Geschwindigkeit und Genauigkeit, was es ideal für Produktionsanwendungen macht. Für das Training verwenden wir die Bibliothek super-gradients, die von den Entwicklern von YOLO-NAS bereitgestellt wird.

Schritt für Schritt – how to train yolo on custom dataset

Beginnen wir mit dem Tutorial:

Abhängigkeitsinstallation für die Objekterkennung

Für das Training und Testen des Modells sind folgende Bibliotheken erforderlich:

pip install -q super-gradients
pip install -q pycocotools
pip install -q onnx
pip install -q onnxruntime

Durch den Export des Modells in das ONNX-Format kann das Modell unabhängig von der Trainingsbibliothek eingesetzt werden.

Notwendige Importe für die Objekterkennung

# General imports
from pathlib import Path
import datetime
# Data imports
from super_gradients.training.datasets.detection_datasets.coco_format_detection import COCOFormatDetectionDataset
from super_gradients.training.utils.collate_fn.crowd_detection_collate_fn import CrowdDetectionCollateFN
from super_gradients.training.transforms.transforms import DetectionMosaic, DetectionRandomAffine, DetectionHSV, DetectionHorizontalFlip, DetectionVerticalFlip, DetectionPaddedRescale, DetectionStandardize, DetectionTargetsFormatTransform
from super_gradients.training import dataloaders
from super_gradients.training.datasets.datasets_utils import worker_init_reset_seed
# Training imports
from super_gradients.training import Trainer
from super_gradients.common.object_names import Models
from super_gradients.training import models
from super_gradients.training.losses import PPYoloELoss
from super_gradients.training.metrics import DetectionMetrics_050
from super_gradients.training.utils.distributed_training_utils import setup_device
from super_gradients.training.models.detection_models.pp_yolo_e import PPYoloEPostPredictionCallback
# Test and export imports
import torch
import torchvision
import onnx
import numpy as np
from PIL import Image
from onnxruntime import InferenceSession

Hyperparameter-Einstellung für das Objekterkennungs-Modell

# experiment name definition
t = datetime.datetime.now()
EXPERIEMENT_NAME = f"{t.year}-{t.month}-{t.tag}-{t.stunde}-{t.minute}-checkbox-detector"
# Model params
MODEL_NAME = "YOLO_NAS_S"
SIZE = 1280
# Training params
WARMUP_INITIAL_LR = 5e-4
INITIAL_LR = 1e-3
COSINE_FINAL_LR_RATIO = 0.01
ZERO_WEIGHT_DECAY_ON_BIAS_AND_BN = True
LR_WARMUP_EPOCHS = 1
OPTIMIZER_WEIGHT_DECAY = 1e-4
EMA = True
EMA_DECAY = 0.9999
MAX_EPOCHS = 20
BATCH_SIZE_TRAIN = 2
BATCH_SIZE_TEST = 6
MIXED_PRECISION = True
# Data params
IGNORE_EMPTY_ANNOTATIONS = True

Dataset und Dataloader für die Objekterkennung

# Base path to dataset
BASE_PATH= Path(".").absolute() # modify accordingly
# Path to specific dataset
TRAIN_FOLDER = "dataset"
TRAIN_ANNOTATION_FILE = "train.json"
TEST_FOLDER = "dataset"
TEST_ANNOTATION_FILE = "val.json"
# train path
train_path = BASE_PATH / TRAIN_FOLDER
train_img_path = train_path /  "images"
train_ann_path = train_path / TRAIN_ANNOTATION_FILE
# test path
test_path = BASE_PATH / TEST_FOLDER
test_img_path = test_path / "images"
test_ann_path = test_path / TEST_ANNOTATION_FILE
# checks
assert train_path.exists(), f"Train path {train_path} does not exist"
assert train_img_path.exists(), f"Train image path {train_img_path} does not exist"
assert train_ann_path.exists(), f"Train annotation path {train_ann_path} does not exist"
assert test_path.exists(), f"Train path {test_path} does not exist"
assert test_img_path.exists(), f"Train image path {test_img_path} does not exist"
assert test_ann_path.exists(), f"Train annotation path {test_ann_path} does not exist"
# train dataset
trainset = COCOFormatDetectionDataset(data_dir=str(train_path),
                                      images_dir="",
                                      json_annotation_file=str(train_ann_path),
                                      input_dim=None,
                                      ignore_empty_annotations=IGNORE_EMPTY_ANNOTATIONS,
                                      transforms=[
                                          DetectionMosaic(prob=1., input_dim=(SIZE, SIZE)),
                                          DetectionRandomAffine(degrees=0.5, scales=(0.9, 1.1), shear=0.0,
                                                                target_size=(SIZE, SIZE),
                                                                filter_box_candidates=False, border_value=114),
                                          DetectionHSV(prob=1., hgain=1, vgain=6, sgain=6),
                                          DetectionHorizontalFlip(prob=0.5),
                                          DetectionVerticalFlip(prob=0.5),
                                          DetectionPaddedRescale(input_dim=(SIZE, SIZE)),
                                          DetectionStandardize(max_value=255),
                                          DetectionTargetsFormatTransform(input_dim=(SIZE, SIZE), output_format="LABEL_CXCYWH")
                                      ])
# validation dataset
valset = COCOFormatDetectionDataset(data_dir=str(test_path),
                                    images_dir="",
                                    json_annotation_file=str(test_ann_path),
                                    input_dim=None,
                                    ignore_empty_annotations=False,
                                    transforms=[
                                        DetectionPaddedRescale(input_dim=(SIZE, SIZE)),
                                        DetectionStandardize(max_value=255),
                                        DetectionTargetsFormatTransform(input_dim=(SIZE, SIZE), output_format="LABEL_CXCYWH")
                                    ])
# get number of classes from dataset
num_classes = len(trainset.classes)
# train dataloader
train_loader = dataloaders.get(dataset=trainset, dataloader_params={
    "shuffle": True,
    "batch_size": BATCH_SIZE_TRAIN,
    "drop_last": False,
    "pin_memory": True,
    "collate_fn": CrowdDetectionCollateFN(),
    "worker_init_fn": worker_init_reset_seed,
    "min_samples": 512
})
# validation dataloader
valid_loader = dataloaders.get(dataset=valset, dataloader_params={
    "shuffle": False,
    "batch_size": BATCH_SIZE_TEST,
    "num_workers": 2,
    "drop_last": False,
    "pin_memory": True,
    "collate_fn": CrowdDetectionCollateFN(),
    "worker_init_fn": worker_init_reset_seed
})

Training des Objekterkennungs-Modells

# training parameter definition for trainer
train_params = {
  "warmup_initial_lr": WARMUP_INITIAL_LR,
  "initial_lr": INITIAL_LR,
  "lr_mode": "cosine",
  "cosine_final_lr_ratio": COSINE_FINAL_LR_RATIO,
  "optimizer": "AdamW",
  "zero_weight_decay_on_bias_and_bn": ZERO_WEIGHT_DECAY_ON_BIAS_AND_BN,
  "lr_warmup_epochs": LR_WARMUP_EPOCHS,
  "warmup_mode": "linear_epoch_step",
  "optimizer_params": {"weight_decay": OPTIMIZER_WEIGHT_DECAY},
  "ema": EMA,
  "ema_params": {"decay": EMA_DECAY, "decay_type": "threshold"},
  "max_epochs": MAX_EPOCHS,
  "mixed_precision": MIXED_PRECISION,
  "loss": PPYoloELoss(use_static_assigner=False, num_classes=num_classes, reg_max=16),
  "valid_metrics_list": [
      DetectionMetrics_050(score_thres=0.1, num_cls=num_classes, normalize_targets=True, post_prediction_callback=PPYoloEPostPredictionCallback(score_threshold=0.01, nms_top_k=1000, max_predictions=300, nms_threshold=0.7))],
  "metric_to_watch": 'F1@0.50',
  }
# model selection
if MODEL_NAME=="YOLO_NAS_S":
  model = Models.YOLO_NAS_S
elif MODEL_NAME=="YOLO_NAS_M":
  model = Models.YOLO_NAS_M
elif MODEL_NAME=="YOLO_NAS_L":
  model = Models.YOLO_NAS_L
# device selection
setup_device(device="cuda")
# define trainer
trainer = Trainer(experiment_name=EXPERIEMENT_NAME, ckpt_root_dir="./checkpoints_dir")
# define model
yolo_model = models.get(model, num_classes=num_classes, pretrained_weights=None)
# start training
trainer.train(model=yolo_model, training_params=train_params, train_loader=train_loader, valid_loader=valid_loader)

Export und Speicherung des Objekterkennungs-Modells

# export parameters
BATCH_SIZE = 1
CHANNELS = 3
WEIGHTS_FILE = "ckpt_best.pth"
EXPORT_NAME = "yolo_model.onnx"
# get path to trained model
checkpoint_dir = Path(trainer.sg_logger._local_dir).absolute()
checkpoint_path = checkpoint_dir / WEIGHTS_FILE
assert checkpoint_path.exists(), f"No checkpoint file found in {checkpoint_path}. Check if the train run was successful."
# load the trained model
yolo_model = models.get(model, num_classes=num_classes, checkpoint_path=str(checkpoint_path))
yolo_model.to(trainer.device)
# define dummy input
dummy_input = torch.randn(BATCH_SIZE, CHANNELS, SIZE, SIZE, device=trainer.device)
# define input and output names
input_names = ["input"]
output_names = ["output"]
# export the model to onnx format
torch.onnx.export(yolo_model, dummy_input, EXPORT_NAME, verbose=False, input_names=input_names, output_names=output_names)
assert Path(EXPORT_NAME).exists(), "\nModel export was not successful.\n"
# check the onnx model
model_onnx = onnx.load(EXPORT_NAME)
onnx.checker.check_model(model_onnx)
print("\nModel exported to ONNX format.\n")

Laden und Testen des Objekterkennungs-Modells

class Detector:
    """Detect checkboxes in images using a pre-trained model."""
    def __init__(self, onnx_path, input_shape, num_classes, threshold=0.7):
        """Initialize the CheckboxDetector with a pre-trained model and default parameter."""
        self.session = InferenceSession(onnx_path)
        self.input_shape = input_shape
        self.threshold = threshold
        self.num_classes = num_classes  
    def __call__(self, image):
        """Run model inference and pre/post processing."""
        input_image = self._preprocess(image)
        outputs = self.session.run(None, {"input": input_image})
        cls_conf, bboxes = self._postprocess(outputs, image.size)
        return cls_conf, bboxes
    def _threshold(self, cls_conf, bboxes):
        """Filter detections based on confidence threshold."""
        idx = np.argwhere(cls_conf > self.threshold)
        cls_conf = cls_conf[idx[:, 0]]
        bboxes = bboxes[idx[:, 0]]
        return cls_conf, bboxes
    def _nms(self, cls_conf, bboxes):
        """Apply Non-Maximum-Suppression to detections."""
        indices = torchvision.ops.nms(torch.from_numpy(bboxes), torch.from_numpy(cls_conf.max(1)), iou_threshold=0.5).numpy()
        cls_conf = cls_conf[indices]
        bboxes = bboxes[indices]
        return cls_conf, bboxes
    def _rescale(self, image, output_shape):
        """Rescale image to a specified output shape."""
        height, width = image.shape[:2]
        scale_factor = min(output_shape[0] / height, output_shape[1] / width)
        if scale_factor != 1.0:
            new_height, new_width = (round(height * scale_factor), round(width * scale_factor))
            image = Image.fromarray(image)
            image = image.resize((new_width, new_height), Image.LANCZOS)
            image = np.array(image)
        return image
    def _bottom_right_pad(self, image, output_shape, pad_value = (114, 114, 114)):
        """Pad image on the bottom and right to reach the output shape."""
        height, width = image.shape[:2]
        pad_height = output_shape[0] - height
        pad_width = output_shape[1] - width
        pad_h = (0, pad_height)  # top=0, bottom=pad_height
        pad_w = (0, pad_width)  # left=0, right=pad_width
        constant_values = ((pad_value, pad_value), (pad_value, pad_value), (0, 0))
        constant_values = np.array(constant_values, dtype=np.object_)
        padding_values = (pad_h, pad_w, (0, 0))
        processed_image = np.pad(image, pad_width=padding_values, mode="constant", constant_values=constant_values)
        return processed_image
    def _permute(self, image, permutation = (2, 0, 1)):
        """Permute the image channels."""
        processed_image = np.ascontiguousarray(image.transpose(permutation))
        return processed_image
    def _standardize(self, image, max_value=255):
        """Standardize the pixel values of image."""
        processed_image = (image / max_value).astype(np.float32)
        return processed_image
    def _preprocess(self, image):
        """Preprocesses image with all transforms as during training before passing it to the model."""
        if image.mode == "P":
            image = image.convert("RGB")
        image = np.array(image)[:, :, ::-1]  # convert to np and BGR as during training
        image = self._rescale(image, output_shape=self.input_shape)
        image = self._bottom_right_pad(image, output_shape=self.input_shape, pad_value=(114, 114, 114))
        image = self._permute(image, permutation=(2, 0, 1))
        image = self._standardize(image, max_value=255)
        image = image[np.newaxis, ...]  # add batch dimension
        return image
    def _postprocess(self, outputs, image_shape):
        """Postprocesses the model's outputs to obtain final detections."""
        bboxes = outputs[0][0,:,:]
        cls_conf = outputs[1][0,:,:]
        cls_conf, bboxes = self._threshold(cls_conf, bboxes)
        if len(cls_conf) > 1:
            cls_conf, bboxes = self._nms(cls_conf, bboxes)
        #Define and apply scale for the bounding boxes to the original image size
        scaler = max((image_shape[1] / self.input_shape[1], image_shape[0] / self.input_shape[0]))
        bboxes *= scaler
        bboxes = np.array([(int(b[0]), int(b[1]), int(b[2]), int(b[3])) for b in bboxes])
        return cls_conf, bboxes
# instantiate detector
detector = Detector(onnx_path=EXPORT_NAME, input_shape=(SIZE,SIZE), num_classes=num_classes, threshold=0.7)
# load test image
sample_img_path = Path("./test_image.png")
assert sample_img_path.exists(), f"Image file with path {sample_img_path} not found."
sample_img = Image.open(str(sample_img_path), mode='r')
# run inference
cls_conf, bboxes = detector(sample_img)
checked = [True if c[0] > c[1] else False for c in cls_conf]
score = cls_conf.max(1)
# vizualization
import matplotlib.pyplot as plt
import copy
%matplotlib inline
import matplotlib as mpl
from PIL import Image, ImageDraw
mpl.rcParams['figure.dpi']= 600
colors = [(0,1,0), (0,1,1)]
colorst = [(1,1,0), (1,0,1)]
def plot_results(pil_img, scores, labels, boxes, name=None):
    plt.figure(figsize=(2,1))
    fig, ax = plt.subplots(1,2)
    ax[0].axis('off')
    ax[0].imshow(copy.deepcopy(pil_img))
    for score, label, (xmin, ymin, xmax, ymax) in zip(scores, labels, boxes):
        ax[1].add_patch(plt.Rectangle((xmin, ymin), xmax - xmin, ymax - ymin, fill=False, color=colors[int(label)], linewidth=1))
        text = f'{score:0.2f}'
        ax[1].text(xmin, ymin, text, fontsize=2, bbox=dict(alpha=0.0))
    ax[1].axis('off')
    ax[1].imshow(pil_img)
    fig.show()
    fig.savefig(f'{name}')
# show result
plot_results(sample_img, score, checked, bboxes, "Example of checkbox detection")

Fazit zum Tutorial

In diesem Tutorial haben wir gezeigt, wie man ein Modell zur Objekterkennung, speziell YOLO-NAS, auf einem COCO-Dataset trainiert und testet. Objekterkennung ist eine Schlüsseltechnologie, die in vielen verschiedenen Bereichen eingesetzt werden kann, um wertvolle Informationen aus Bildern zu extrahieren.

Anwendungsbeispiele von finden Sie in unserem informativen Artikel auf dem Konfuzio Blog: YOLO NAS: Object Detection Model

Weitere technische Details und Hintergrundinformationen finden Sie in der ausführlichen Konfuzio Dokumentation: dev.konfuzio.com

Objekte und deren Merkmale können in Bildern erkannt und weiterverarbeitet werden, was in der Bildverarbeitung, der Überwachung und vielen anderen Bereichen nützlich ist. Durch den Einsatz von Deep Learning und anderen Techniken können wir effizienter arbeiten und bessere Produkte entwickeln.

Die Fähigkeit, Objekte in Bildern, Scans und im Allgemeinen digitalisierten Dokumenten genau zu erkennen, eröffnet viele Möglichkeiten für zukünftige Projekte und Anwendungen.

Fanden Sie die Seite hilfreich?

Vielen Dank für Ihr Feedback!

Würden Sie mir Feedback geben? (anonym)

Wir entwickeln KI-Software für Unternehmen und verzichten bewusst auf störende Werbebanner. Durch unsere Artikel dokumentieren wir Themen, die uns beschäftigen, interessieren und auch unser tägliches Brot finanzieren.

Da unsere Inhalte kostenfrei sind, ist Ihr Feedback unser Lob.

Jeder Autor liest Ihr anonymes Feedback persönlich, obwohl KI es automatisieren könnten, und integriert konstruktive Vorschläge direkt in die nächste Überarbeitung oder nutzt es als Inspiration für den nächsten Artikel.



    • Nico Engelmann
      (Autor)

      Als erfahrener AI Engineer lasse ich Lösungen mittels Künstlicher Intelligenz und klassischen Algorithmen entstehen.

    de_DEDE