GPU & Accelerated Computing Bundle | Prompeteer.ai
CUDA and GPU-accelerated data science (RAPIDS/cuDF, cuPyNumeric), HPC and Slurm workflows, distributed LLM training & inference (NeMo, Megatron), numerical optimization (cuOpt), and quantum computing (CUDA-Q) for accelerated-computing engineers.
Included Skills (89)
- NeMo Distributed Trainer — This tool assists developers in configuring and optimizing distributed training strategies within NeMo AutoModel to ensure efficient multi-GPU and multi-node model performance.
- Memory Optimization Assistant — This tool helps AI engineers optimize GPU memory usage by configuring selective or full activation recompute strategies within the Megatron Bridge framework.
- Parallelism Strategy Optimizer — This guide assists AI engineers in selecting and configuring optimal parallelism strategies for Megatron Bridge to maximize performance and hardware utilization.
- Megatron Resiliency Manager — This tool enables and manages fault tolerance, straggler detection, and automatic restart features to ensure stable training for Megatron Bridge users.
- GPU Portfolio Optimizer — This tool enables quantitative analysts to build, backtest, and optimize high-performance financial portfolios using NVIDIA-accelerated Mean-CVaR algorithms and GPU-based solvers.
- Megatron CPU Offloading — This skill helps developers optimize GPU memory usage in Megatron Bridge by configuring activation and optimizer state offloading to CPU memory.
- MoE Parallel Optimizer — This skill helps AI engineers optimize Megatron-Bridge performance by configuring expert-parallel communication overlap and dispatcher backends to reduce training latency.
- TAO Execution SDK — This SDK enables developers to manage, monitor, and scale NVIDIA TAO GPU training jobs across diverse cloud and local computing platforms.
- NeMo Job Launcher — This tool helps developers configure NeMo AutoModel job submissions for interactive environments, Slurm HPC clusters, and SkyPilot cloud-based execution platforms.
- TAO Docker Executor — This tool enables developers to execute TAO SDK jobs within local or remote Docker containers utilizing NVIDIA GPU acceleration for efficient model training.
- DALI Dynamic Assistant — This assistant helps developers write, review, and migrate deep learning data pipelines using the NVIDIA DALI imperative dynamic-mode API for GPU processing.
- Optimization Formulation Guide — This guide assists developers in translating complex business problems into structured LP, MILP, or QP mathematical models for efficient numerical optimization.
- Distributed Data Loader — This skill enables developers to efficiently load complex, sharded datasets into distributed cuPyNumeric ndarrays when standard built-in loaders are insufficient for custom layouts.
- CuTile Triton Converter — This tool automates the conversion and debugging of cuTile GPU kernels into Triton, assisting developers in porting high-performance code with optimized mapping.
- SLURM Cluster Executor — This tool enables AI agents to execute TAO training and inference jobs on remote SLURM GPU clusters using SSH, Pyxis, and Enroot containers.
- cuTile Kernel Assistant — This expert assistant helps developers write, debug, and optimize high-performance GPU kernels using the cuTile programming model for NVIDIA hardware.
- Brev GPU Orchestrator — This tool enables developers to manage NVIDIA TAO training and inference workloads by automating GPU instance deployment and job execution via the Brev CLI.
- Clinical ASR Builder — This tool assists clinical researchers in curating medical terminology, generating IPA tags, and synthesizing NeMo manifests for ASR model development.
- Clinical ASR Evaluator — This tool scores NeMo manifests to generate comprehensive KER leaderboards, assisting clinical AI researchers in evaluating and benchmarking speech recognition performance.
- Clinical ASR Bootstrapper — This tool automates the initial configuration and smoke testing of clinical ASR environments to help developers quickly bootstrap healthcare AI workflows.
- NeMo Model Onboarding — This guide assists developers in integrating new model architectures into NeMo AutoModel through structured implementation, registration, and validation workflows.
- NeMo Recipe Architect — This tool assists developers in creating, configuring, and validating NeMo AutoModel training and evaluation recipes through structured YAML and builder management.
- TileGym Kernel Integrator — This tool guides developers through the end-to-end process of registering, implementing, and benchmarking new cuTile GPU operators within the TileGym framework.
- HSB FPGA Flasher — This tool enables developers to safely flash FPGA firmware on HSB Lattice boards and Leopard Imaging VB1940 cameras connected to NVIDIA devkits.
- Holoscan Bridge Setup — This tool automates the end-to-end deployment, container configuration, and connectivity verification of the NVIDIA Holoscan Sensor Bridge for developers and engineers.
- Megatron Testing Framework — This tool assists developers in managing unit and functional tests, configuring recipe YAML files, and maintaining golden values for the Megatron-LM ecosystem.
- Sandbox Network Policy — Configures allowed egress endpoints for NemoClaw sandboxes to help developers manage network access rules and security policies for their applications.
- Lepton GPU Orchestrator — This tool enables developers to manage, dispatch, and monitor TAO container workloads on DGX Cloud Lepton GPU infrastructure without manual cluster configuration.
- CuTile Autotuning Assistant — This tool assists developers in designing, implementing, and optimizing CuTile autotuning kernels to achieve peak performance across diverse GPU architectures.
- Jetson Clock Customizer — This tool enables developers to configure Jetson CPU, GPU, and EMC clock behaviors by modifying pre-flash system files for optimized hardware performance.
- GPU Memory Optimizer — This skill provides techniques for reducing peak GPU memory usage and resolving OOM errors in Megatron Bridge for deep learning engineers.
- CUDA Graph Optimizer — This tool helps developers optimize training performance by configuring and validating CUDA graph implementations within Megatron Bridge for various model architectures.
- Megatron Bridge Recommender — This tool helps developers select and customize optimal Megatron Bridge training recipes based on specific model architectures, hardware configurations, and training objectives.
- GPU Dataframe Expert — Provides expert guidance for developers accelerating pandas workflows using NVIDIA cuDF and dask-cuDF to achieve high-performance GPU-based data processing and ETL.
- CuOpt Development Assistant — This tool assists developers in modifying, building, testing, and debugging the NVIDIA cuOpt codebase while strictly adhering to established project conventions.
- TAO Kubernetes Executor — This tool enables developers to deploy and manage NVIDIA TAO container jobs as scalable Kubernetes workloads with automated GPU resource scheduling.
- TAO AutoML Optimizer — Automates hyperparameter optimization for NVIDIA TAO models using advanced search algorithms to streamline training workflows and improve model performance for developers.
- VCN Gap Analyzer — This tool identifies weak classification samples for NVIDIA TAO VCN experiments to help engineers optimize decision thresholds and target data for augmentation.
- Nemotron Pipeline Orchestrator — This tool enables developers to plan, configure, and execute end-to-end Nemotron customization workflows, including training, optimization, and evaluation pipelines.
- Nemotron Safety Architect — This tool helps developers create custom safety policies, taxonomy, and inference prompts for NVIDIA Nemotron content-safety guardrails to ensure robust model governance.
- GPU Acceleration Expert — This skill helps developers transform CPU-bound Python code into high-performance GPU-accelerated applications using the NVIDIA RAPIDS ecosystem and CUDA-based libraries.
- Megatron-LM Slurm Orchestrator — This guide assists AI engineers in configuring and launching distributed Megatron-LM training jobs on SLURM clusters with optimized environment and network settings.
- Hierarchical Context Parallelism — This guide assists engineers in configuring and verifying hierarchical context parallelism within Megatron-Bridge to optimize large-scale model training performance.
- Megatron FSDP Optimizer — This guide assists engineers in configuring and verifying Megatron FSDP within Megatron-Bridge to optimize training performance and resolve memory issues.
- Dynamo Interconnect Validator — This tool validates RDMA and NVLink connectivity for Dynamo deployments to ensure reliable disaggregated serving performance for infrastructure engineers and operators.
- DeepStream Pipeline Architect — This skill assists developers in building and optimizing NVIDIA DeepStream 9.0 video analytics pipelines using Python and GStreamer-based inference integration.
- MoE Dispatcher Selector — This tool helps AI engineers select the optimal MoE token dispatcher based on hardware, EP degree, and specific model-family performance patterns.
- CuPyNumeric Migration Assessor — This tool evaluates NumPy codebases for cuPyNumeric compatibility, providing developers with actionable refactoring guidance and migration readiness verdicts for distributed GPU scaling.
- Jetson Inference Optimizer — This tool recommends optimal inference runtimes and memory configuration flags for LLM and VLM workloads running on NVIDIA Jetson hardware platforms.
- PhysicsNeMo Navigator — This tool helps researchers and developers navigate the PhysicsNeMo repository by identifying relevant models, datapipes, and examples for specific SciML and AI4Science tasks.
- Molecular Modeling Assistant — This tool provides a cloud-native Python interface for medicinal chemists to automate complex molecular simulations, drug discovery workflows, and protein-ligand modeling tasks.
- Holoscan Conda Installer — This tool automates the installation of Holoscan SDK v4.3+ into Conda environments for developers working on Linux x86_64 systems with CUDA 13.
- Nemotron Retrieval Assistant — This tool assists developers in planning, debugging, and deploying Nemotron embedding and reranking recipes by managing configurations, execution commands, and project workflows.
- CuOpt Routing Optimizer — This tool enables developers to solve complex vehicle routing problems efficiently using the NVIDIA cuOpt Python API for optimized logistics operations.
- Synthetic Dataset Generator — This tool helps data scientists and developers efficiently create custom synthetic datasets or build automated data generation pipelines for their specific projects.
- Autonomous RL Researcher — This agent automates iterative NeMo-RL experiment lifecycles, helping researchers conduct hypothesis testing, launch reproducible training runs, and track results within a git-based ledger.
- Session Memory Manager — This tool enables coding agents to persist and recover project context across disconnects, restarts, or handoffs, ensuring seamless continuity for complex, long-running development tasks.
- CuTile Kernel Converter — This tool automates the translation of Python cuTile GPU kernels into Julia cuTile.jl code, assisting developers in porting, debugging, and optimizing high-performance GPU applications.
- CuTile Kernel Optimizer — This tool helps developers systematically profile, diagnose, and iteratively tune cuTile GPU kernels to achieve peak performance within the TileGym framework.
- TAO Dataset Validator — This tool executes the tao-daft validate command to help developers verify the structure, schema, and integrity of NVIDIA TAO DAFT datasets.
- Jetson Health Monitor — This tool provides a comprehensive, read-only system snapshot to help developers and AI agents troubleshoot performance, thermal, and resource utilization on Jetson devices.
- Quantum Circuit Developer — This skill enables developers to design, simulate, and optimize NISQ quantum circuits specifically for Google Quantum AI hardware and various partner backends.
- Fluid Dynamics Simulator — This high-performance Python framework enables researchers and engineers to execute complex computational fluid dynamics simulations using advanced pseudospectral methods and parallel computing.
- Computational Resource Auditor — This tool assesses system hardware capabilities to provide strategic optimization recommendations for developers performing intensive scientific computing and large-scale data processing tasks.
- Serverless Cloud Orchestrator — This tool enables developers to deploy and scale Python applications, AI models, and GPU-accelerated workloads on a serverless cloud infrastructure.
- Polars Data Processor — This skill provides high-performance DataFrame manipulation and lazy evaluation capabilities to help data engineers and analysts optimize complex ETL and analytics pipelines.
- Quantum Circuit Framework — This framework enables developers to build, optimize, and execute quantum circuits across various hardware platforms, streamlining enterprise-grade quantum computing and algorithm development.
- Time Series Forecaster — This skill provides zero-shot univariate time series forecasting using Google's TimesFM model, enabling data analysts to generate accurate predictions without custom training.
- PEFT Fine-Tuning — Fine-tune LLMs efficiently using LoRA, QLoRA, and other PEFT methods, helping users adapt models to specific tasks on consumer GPUs.
- Vision Pipeline Specialist — Assists users with object detection, image segmentation, and deploying optimized computer vision pipelines using YOLO, SAM, and TensorRT.
- Azure Batch Automation — Automates Azure Batch job management for Java developers, simplifying large-scale parallel and HPC workload execution.
- TAO Dataset Converter — This tool automates the conversion of NVIDIA TAO DAFT datasets between supported formats to assist developers in managing their data pipelines efficiently.
- Quantum System Simulator — This tool provides comprehensive Python-based simulation capabilities for open and closed quantum systems, assisting physics researchers with dynamics, decoherence, and optical modeling.
- Modal Serverless GPU — Deploys and runs Python functions on cloud GPUs, enabling ML model deployment and inference without infrastructure management for developers.
- Triton Inference Server — Deploys AI models at scale, supporting multiple frameworks and hardware, benefiting data scientists and machine learning engineers.
- MoE Performance Optimizer — This workflow guides engineers through systematic MoE training optimization using the Three Walls framework to maximize throughput and resolve performance regressions.
- Megatron Training Orchestrator — This tool assists developers in executing and correlating Megatron-LM and Megatron Bridge training runs to ensure consistent loss curves and configuration parity.
- MoE Communication Optimizer — This tool helps performance engineers tune expert-parallel communication overlap in Megatron Bridge to maximize throughput for large-scale MoE model training.
- MoE Long-Context Optimizer — This guide provides performance optimization strategies for training long-context MoE models, helping engineers resolve memory constraints and maximize throughput in Megatron Bridge.
- Sequence Packing Optimizer — This skill assists developers in configuring and validating packed sequences and long-context training within Megatron-Bridge to optimize performance for LLM and VLM workloads.
- MoE VLM Trainer — This skill provides expert guidance for training MoE vision language models using Megatron Bridge, helping engineers optimize FSDP and 3D-parallel performance strategies.
- Communication Overlap Optimizer — This guide assists engineers in configuring and verifying TP, DP, and PP communication overlap settings within Megatron-Bridge to maximize training throughput.
- MoE Hardware Optimizer — This tool provides optimized training playbooks and throughput configurations for MoE models, assisting engineers in tuning performance across diverse hardware platforms.
- VCN Image Miner — This tool automates the DEFT embedding and mining workflow to help computer vision engineers augment training datasets with relevant nearest-neighbor source images.
- CuOpt Optimization CLI — This tool enables developers to solve linear, mixed-integer, and quadratic programming problems efficiently using the cuOpt command-line interface with MPS files.
- Nextflow Pipeline Engineer — This skill assists bioinformaticians and data scientists in building, debugging, and scaling reproducible data pipelines using Nextflow and the nf-core framework.
- PyTorch Lightning Assistant — This skill helps deep learning researchers and engineers organize PyTorch code, automate training workflows, and scale neural network models across multi-GPU or TPU environments.
- NemoClaw Sandbox Manager — Deploys and manages NemoClaw, NVIDIA's open-source sandbox, to securely run OpenClaw agents with policy-enforced network, filesystem, and inference controls.
- Pacsomatic Workflow Assistant — This toolkit assists bioinformaticians in validating inputs, generating samplesheets, and managing reproducible execution for the nf-core/pacsomatic tumor-normal analysis pipeline.