Lancement / Blog / Instructions / YOLO-NAS Finetuning avec ses propres annotations OMR

YOLO-NAS Finetuning avec ses propres annotations OMR

Résumer avec ChatGPT

La reconnaissance d'objets est l'une des applications les plus passionnantes et les plus polyvalentes de l'intelligence artificielle. Du traitement d'images à la surveillance de la sécurité en passant par les véhicules autonomes, la capacité à reconnaître avec précision des objets dans des images offre de nombreuses possibilités et avantages.

Dans cet article, nous vous montrons comment utiliser la dernière technologie en matière de reconnaissance d'objets, le Modèle de NAS YOLOLes utilisateurs doivent pouvoir s'entraîner et utiliser un modèle de reconnaissance performant.

Objectif du tutoriel

Ce guide montre comment créer un Modèle OMR sur données annotées s'entraîner au format COCO.

Nous allons utiliser le modèle de Reconnaissance des cases à cocher dans les structures de documents mais vous pouvez également adapter le modèle à d'autres cas d'utilisation.

Conditions préalables à la reconnaissance d'objets

  • Compréhension générale de la reconnaissance d'objets et de l'apprentissage des réseaux neuronaux
  • Un jeu de données de reconnaissance d'objets formaté COCO
  • Environnement de développement de modèles de reconnaissance d'objets

Pour démarrer rapidement, vous pouvez créer un nouveau notebook Google Colab ou exécuter le code localement ou dans l'environnement de votre choix.

Aperçu du modèle YOLO-NAS pour la reconnaissance d'objets

Nous utilisons le modèle de reconnaissance d'objets ultramoderne YOLO-NAS. Ce modèle combine vitesse élevée et précision, ce qui le rend idéal pour les applications de production. Pour l'entraînement, nous utilisons la bibliothèque super-gradients, fournie par les développeurs de YOLO-NAS.

Pas à pas - how to train yolo on custom dataset

Commençons par le tutoriel :

Installation de la dépendance pour la reconnaissance des objets

Les bibliothèques suivantes sont nécessaires pour l'entraînement et les tests du modèle :

pip install -q super-gradients
pip install -q pycocotools
pip install -q onnx
pip install -q onnxruntime

L'exportation du modèle au format ONNX permet d'utiliser le modèle indépendamment de la bibliothèque de formation.

Importations nécessaires pour la reconnaissance des objets

# General imports
from pathlib import Path
import datetime
# Data imports
from super_gradients.training.datasets.detection_datasets.coco_format_detection import COCOFormatDetectionDataset
from super_gradients.training.utils.collate_fn.crowd_detection_collate_fn import CrowdDetectionCollateFN
from super_gradients.training.transforms.transforms import DetectionMosaic, DetectionRandomAffine, DetectionHSV, DetectionHorizontalFlip, DetectionVerticalFlip, DetectionPaddedRescale, DetectionStandardize, DetectionTargetsFormatTransform
from super_gradients.training import dataloaders
from super_gradients.training.datasets.datasets_utils import worker_init_reset_seed
# Training imports
from super_gradients.training import Trainer
from super_gradients.common.object_names import Models
from super_gradients.training import models
from super_gradients.training.losses import PPYoloELoss
from super_gradients.training.metrics import DetectionMetrics_050
from super_gradients.training.utils.distributed_training_utils import setup_device
from super_gradients.training.models.detection_models.pp_yolo_e import PPYoloEPostPredictionCallback
# Test and export imports
import torch
import torchvision
import onnx
import numpy as np
from PIL import Image
from onnxruntime import InferenceSession

Réglage des hyperparamètres pour le modèle de reconnaissance d'objets

# experiment name definition
t = datetime.datetime.now()
EXPERIEMENT_NAME = f"{t.year}-{t.month}-{t.tag}-{t.stunde}-{t.minute}-checkbox-detector"
# Model params
MODEL_NAME = "YOLO_NAS_S"
SIZE = 1280
# Training params
WARMUP_INITIAL_LR = 5e-4
INITIAL_LR = 1e-3
COSINE_FINAL_LR_RATIO = 0.01
ZERO_WEIGHT_DECAY_ON_BIAS_AND_BN = True
LR_WARMUP_EPOCHS = 1
OPTIMIZER_WEIGHT_DECAY = 1e-4
EMA = True
EMA_DECAY = 0.9999
MAX_EPOCHS = 20
BATCH_SIZE_TRAIN = 2
BATCH_SIZE_TEST = 6
MIXED_PRECISION = True
# Data params
IGNORE_EMPTY_ANNOTATIONS = True

Dataset et Dataloader pour la reconnaissance d'objets

# Base path to dataset
BASE_PATH= Path(".").absolute() # modify accordingly
# Path to specific dataset
TRAIN_FOLDER = "dataset"
TRAIN_ANNOTATION_FILE = "train.json"
TEST_FOLDER = "dataset"
TEST_ANNOTATION_FILE = "val.json"
# train path
train_path = BASE_PATH / TRAIN_FOLDER
train_img_path = train_path /  "images"
train_ann_path = train_path / TRAIN_ANNOTATION_FILE
# test path
test_path = BASE_PATH / TEST_FOLDER
test_img_path = test_path / "images"
test_ann_path = test_path / TEST_ANNOTATION_FILE
# checks
assert train_path.exists(), f"Train path {train_path} does not exist"
assert train_img_path.exists(), f"Train image path {train_img_path} does not exist"
assert train_ann_path.exists(), f"Train annotation path {train_ann_path} does not exist"
assert test_path.exists(), f"Train path {test_path} does not exist"
assert test_img_path.exists(), f"Train image path {test_img_path} does not exist"
assert test_ann_path.exists(), f"Train annotation path {test_ann_path} does not exist"
# train dataset
trainset = COCOFormatDetectionDataset(data_dir=str(train_path),
                                      images_dir="",
                                      json_annotation_file=str(train_ann_path),
                                      input_dim=None,
                                      ignore_empty_annotations=IGNORE_EMPTY_ANNOTATIONS,
                                      transforms=[
                                          DetectionMosaic(prob=1., input_dim=(SIZE, SIZE)),
                                          DetectionRandomAffine(degrees=0.5, scales=(0.9, 1.1), shear=0.0,
                                                                target_size=(SIZE, SIZE),
                                                                filter_box_candidates=False, border_value=114),
                                          DetectionHSV(prob=1., hgain=1, vgain=6, sgain=6),
                                          DetectionHorizontalFlip(prob=0.5),
                                          DetectionVerticalFlip(prob=0.5),
                                          DetectionPaddedRescale(input_dim=(SIZE, SIZE)),
                                          DetectionStandardize(max_value=255),
                                          DetectionTargetsFormatTransform(input_dim=(SIZE, SIZE), output_format="LABEL_CXCYWH")
                                      ])
# validation dataset
valset = COCOFormatDetectionDataset(data_dir=str(test_path),
                                    images_dir="",
                                    json_annotation_file=str(test_ann_path),
                                    input_dim=None,
                                    ignore_empty_annotations=False,
                                    transforms=[
                                        DetectionPaddedRescale(input_dim=(SIZE, SIZE)),
                                        DetectionStandardize(max_value=255),
                                        DetectionTargetsFormatTransform(input_dim=(SIZE, SIZE), output_format="LABEL_CXCYWH")
                                    ])
# get number of classes from dataset
num_classes = len(trainset.classes)
# train dataloader
train_loader = dataloaders.get(dataset=trainset, dataloader_params={
    "shuffle": True,
    "batch_size": BATCH_SIZE_TRAIN,
    "drop_last": False,
    "pin_memory": True,
    "collate_fn": CrowdDetectionCollateFN(),
    "worker_init_fn": worker_init_reset_seed,
    "min_samples": 512
})
# validation dataloader
valid_loader = dataloaders.get(dataset=valset, dataloader_params={
    "shuffle": False,
    "batch_size": BATCH_SIZE_TEST,
    "num_workers": 2,
    "drop_last": False,
    "pin_memory": True,
    "collate_fn": CrowdDetectionCollateFN(),
    "worker_init_fn": worker_init_reset_seed
})

Entraînement au modèle de reconnaissance d'objets

# training parameter definition for trainer
train_params = {
  "warmup_initial_lr": WARMUP_INITIAL_LR,
  "initial_lr": INITIAL_LR,
  "lr_mode": "cosine",
  "cosine_final_lr_ratio": COSINE_FINAL_LR_RATIO,
  "optimizer": "AdamW",
  "zero_weight_decay_on_bias_and_bn": ZERO_WEIGHT_DECAY_ON_BIAS_AND_BN,
  "lr_warmup_epochs": LR_WARMUP_EPOCHS,
  "warmup_mode": "linear_epoch_step",
  "optimizer_params": {"weight_decay": OPTIMIZER_WEIGHT_DECAY},
  "ema": EMA,
  "ema_params": {"decay": EMA_DECAY, "decay_type": "threshold"},
  "max_epochs": MAX_EPOCHS,
  "mixed_precision": MIXED_PRECISION,
  "loss": PPYoloELoss(use_static_assigner=False, num_classes=num_classes, reg_max=16),
  "valid_metrics_list": [
      DetectionMetrics_050(score_thres=0.1, num_cls=num_classes, normalize_targets=True, post_prediction_callback=PPYoloEPostPredictionCallback(score_threshold=0.01, nms_top_k=1000, max_predictions=300, nms_threshold=0.7))],
  "metric_to_watch": 'F1@0.50',
  }
# model selection
if MODEL_NAME=="YOLO_NAS_S":
  model = Models.YOLO_NAS_S
elif MODEL_NAME=="YOLO_NAS_M":
  model = Models.YOLO_NAS_M
elif MODEL_NAME=="YOLO_NAS_L":
  model = Models.YOLO_NAS_L
# device selection
setup_device(device="cuda")
# define trainer
trainer = Trainer(experiment_name=EXPERIEMENT_NAME, ckpt_root_dir="./checkpoints_dir")
# define model
yolo_model = models.get(model, num_classes=num_classes, pretrained_weights=None)
# start training
trainer.train(model=yolo_model, training_params=train_params, train_loader=train_loader, valid_loader=valid_loader)

Exportation et enregistrement du modèle de reconnaissance d'objets

# export parameters
BATCH_SIZE = 1
CHANNELS = 3
WEIGHTS_FILE = "ckpt_best.pth"
EXPORT_NAME = "yolo_model.onnx"
# get path to trained model
checkpoint_dir = Path(trainer.sg_logger._local_dir).absolute()
checkpoint_path = checkpoint_dir / WEIGHTS_FILE
assert checkpoint_path.exists(), f"No checkpoint file found in {checkpoint_path}. Check if the train run was successful."
# load the trained model
yolo_model = models.get(model, num_classes=num_classes, checkpoint_path=str(checkpoint_path))
yolo_model.to(trainer.device)
# define dummy input
dummy_input = torch.randn(BATCH_SIZE, CHANNELS, SIZE, SIZE, device=trainer.device)
# define input and output names
input_names = ["input"]
output_names = ["output"]
# export the model to onnx format
torch.onnx.export(yolo_model, dummy_input, EXPORT_NAME, verbose=False, input_names=input_names, output_names=output_names)
assert Path(EXPORT_NAME).exists(), "\nModel export was not successful.\n"
# check the onnx model
model_onnx = onnx.load(EXPORT_NAME)
onnx.checker.check_model(model_onnx)
print("\nModel exported to ONNX format.\n")

Chargement et test du modèle de reconnaissance d'objets

class Detector:
    """Detect checkboxes in images using a pre-trained model."""
    def __init__(self, onnx_path, input_shape, num_classes, threshold=0.7):
        """Initialize the CheckboxDetector with a pre-trained model and default parameter."""
        self.session = InferenceSession(onnx_path)
        self.input_shape = input_shape
        self.threshold = threshold
        self.num_classes = num_classes  
    def __call__(self, image):
        """Run model inference and pre/post processing."""
        input_image = self._preprocess(image)
        outputs = self.session.run(None, {"input": input_image})
        cls_conf, bboxes = self._postprocess(outputs, image.size)
        return cls_conf, bboxes
    def _threshold(self, cls_conf, bboxes):
        """Filter detections based on confidence threshold."""
        idx = np.argwhere(cls_conf > self.threshold)
        cls_conf = cls_conf[idx[:, 0]]
        bboxes = bboxes[idx[:, 0]]
        return cls_conf, bboxes
    def _nms(self, cls_conf, bboxes):
        """Apply Non-Maximum-Suppression to detections."""
        indices = torchvision.ops.nms(torch.from_numpy(bboxes), torch.from_numpy(cls_conf.max(1)), iou_threshold=0.5).numpy()
        cls_conf = cls_conf[indices]
        bboxes = bboxes[indices]
        return cls_conf, bboxes
    def _rescale(self, image, output_shape):
        """Rescale image to a specified output shape."""
        height, width = image.shape[:2]
        scale_factor = min(output_shape[0] / height, output_shape[1] / width)
        if scale_factor != 1.0:
            new_height, new_width = (round(height * scale_factor), round(width * scale_factor))
            image = Image.fromarray(image)
            image = image.resize((new_width, new_height), Image.LANCZOS)
            image = np.array(image)
        return image
    def _bottom_right_pad(self, image, output_shape, pad_value = (114, 114, 114)):
        """Pad image on the bottom and right to reach the output shape."""
        height, width = image.shape[:2]
        pad_height = output_shape[0] - height
        pad_width = output_shape[1] - width
        pad_h = (0, pad_height)  # top=0, bottom=pad_height
        pad_w = (0, pad_width)  # left=0, right=pad_width
        constant_values = ((pad_value, pad_value), (pad_value, pad_value), (0, 0))
        constant_values = np.array(constant_values, dtype=np.object_)
        padding_values = (pad_h, pad_w, (0, 0))
        processed_image = np.pad(image, pad_width=padding_values, mode="constant", constant_values=constant_values)
        return processed_image
    def _permute(self, image, permutation = (2, 0, 1)):
        """Permute the image channels."""
        processed_image = np.ascontiguousarray(image.transpose(permutation))
        return processed_image
    def _standardize(self, image, max_value=255):
        """Standardize the pixel values of image."""
        processed_image = (image / max_value).astype(np.float32)
        return processed_image
    def _preprocess(self, image):
        """Preprocesses image with all transforms as during training before passing it to the model."""
        if image.mode == "P":
            image = image.convert("RGB")
        image = np.array(image)[:, :, ::-1]  # convert to np and BGR as during training
        image = self._rescale(image, output_shape=self.input_shape)
        image = self._bottom_right_pad(image, output_shape=self.input_shape, pad_value=(114, 114, 114))
        image = self._permute(image, permutation=(2, 0, 1))
        image = self._standardize(image, max_value=255)
        image = image[np.newaxis, ...]  # add batch dimension
        return image
    def _postprocess(self, outputs, image_shape):
        """Postprocesses the model's outputs to obtain final detections."""
        bboxes = outputs[0][0,:,:]
        cls_conf = outputs[1][0,:,:]
        cls_conf, bboxes = self._threshold(cls_conf, bboxes)
        if len(cls_conf) > 1:
            cls_conf, bboxes = self._nms(cls_conf, bboxes)
        #Define and apply scale for the bounding boxes to the original image size
        scaler = max((image_shape[1] / self.input_shape[1], image_shape[0] / self.input_shape[0]))
        bboxes *= scaler
        bboxes = np.array([(int(b[0]), int(b[1]), int(b[2]), int(b[3])) for b in bboxes])
        return cls_conf, bboxes
# instantiate detector
detector = Detector(onnx_path=EXPORT_NAME, input_shape=(SIZE,SIZE), num_classes=num_classes, threshold=0.7)
# load test image
sample_img_path = Path("./test_image.png")
assert sample_img_path.exists(), f"Image file with path {sample_img_path} not found."
sample_img = Image.open(str(sample_img_path), mode='r')
# run inference
cls_conf, bboxes = detector(sample_img)
checked = [True if c[0] > c[1] else False for c in cls_conf]
score = cls_conf.max(1)
# vizualization
import matplotlib.pyplot as plt
import copy
%matplotlib inline
import matplotlib as mpl
from PIL import Image, ImageDraw
mpl.rcParams['figure.dpi']= 600
colors = [(0,1,0), (0,1,1)]
colorst = [(1,1,0), (1,0,1)]
def plot_results(pil_img, scores, labels, boxes, name=None):
    plt.figure(figsize=(2,1))
    fig, ax = plt.subplots(1,2)
    ax[0].axis('off')
    ax[0].imshow(copy.deepcopy(pil_img))
    for score, label, (xmin, ymin, xmax, ymax) in zip(scores, labels, boxes):
        ax[1].add_patch(plt.Rectangle((xmin, ymin), xmax - xmin, ymax - ymin, fill=False, color=colors[int(label)], linewidth=1))
        text = f'{score:0.2f}'
        ax[1].text(xmin, ymin, text, fontsize=2, bbox=dict(alpha=0.0))
    ax[1].axis('off')
    ax[1].imshow(pil_img)
    fig.show()
    fig.savefig(f'{name}')
# show result
plot_results(sample_img, score, checked, bboxes, "Example of checkbox detection")

Conclusion du tutoriel

Dans ce tutoriel, nous avons montré comment entraîner et tester un modèle de reconnaissance d'objets, en particulier YOLO-NAS, sur un dataset COCO. La reconnaissance d'objets est une technologie clé qui peut être utilisée dans de nombreux domaines différents afin d'extraire des informations précieuses des images.

Vous trouverez des exemples d'application de dans notre article informatif sur le blog Confuzio : YOLO NAS : Modèle de détection d'objets

Vous trouverez plus de détails techniques et d'informations de fond dans la documentation détaillée de Confuzio : dev.konfuzio.com

Les objets et leurs Les caractéristiques peuvent être reconnues et traitées dans les images. ce qui est utile dans le traitement des images, la surveillance et bien d'autres domaines. L'utilisation du deep learning et d'autres techniques nous permet de travailler plus efficacement et de développer de meilleurs produits.

La capacité de reconnaître avec précision des objets dans des images, des scans et des documents numérisés en général ouvre de nombreuses possibilités pour de futurs projets et applications.

Avez-vous trouvé cette page utile ?

Merci beaucoup pour vos commentaires !

Voulez-vous me donner un feedback ? (anonyme)

Nous développons des logiciels d'intelligence artificielle pour les entreprises et renonçons délibérément aux bannières publicitaires gênantes. Par le biais de nos articles, nous documentons des thèmes qui nous préoccupent, nous intéressent et financent également notre pain quotidien.

Comme nos contenus sont gratuits, vos commentaires sont nos félicitations.

Chaque auteur lit personnellement vos commentaires anonymes, même si les IA pourraient les automatiser, et intègre directement les suggestions constructives dans la prochaine révision ou les utilise comme source d'inspiration pour le prochain article.



    • Nico Engelmann
      (auteur)

      En tant qu'ingénieur AI expérimenté, je fais naître des solutions au moyen de l'intelligence artificielle et d'algorithmes classiques.

    fr_FRFR