What Is Automatic Speech Recognition and GPU-Accelerated ASR
Overview of ASR (speech-to-text): pipeline, acoustic and language models, CTC training, decoding strategies, and GPU acceleration using NVIDIA NeMo and toolkits.
Overview of ASR (speech-to-text): pipeline, acoustic and language models, CTC training, decoding strategies, and GPU acceleration using NVIDIA NeMo and toolkits.
LSG-SLAM: a stereo visual SLAM using 3D Gaussian splatting for large-scale outdoor reconstruction, improving tracking stability and mapping quality.
Overview of artificial intelligence and its relationship to machine learning and deep learning, covering AI categories, ML workflow, and common deep architectures.
MAX78000 AI microcontroller with low-power CNN accelerator, Arm Cortex?M4 with FPU, 442 KB weight SRAM and 512 KB flash, optimized for edge inference.
LLM-driven system automates embodied intelligence skill generation for robotic arms, converting natural-language tasks into MuJoCo scenes, actions, and reward code.
Survey of deep metric learning: formulations, sample selection and metric loss functions (contrastive, triplet, N?pair), architectures and applications in vision, audio, and text.
Survey of deep learning approaches for radar target detection, comparing two-stage and single-stage detectors (Faster R-CNN, YOLOv5), preprocessing, and deployment results.
Overview of convolutional neural networks: principles like padding, stride, pooling and filters, edge detection fundamentals, architecture patterns and a Keras MNIST implementation.
Survey of deep learning applications in AI: image recognition, NLP, speech, recommendation systems, autonomous driving, healthcare, cybersecurity, and VR.
Analysis of an IoT smart classroom solution - hardware connectivity, data interoperability and scenario intelligence for unified, energy-efficient device management.