GPU & Accelerated Computing Bundle | Prompeteer.ai

CUDA and GPU-accelerated data science (RAPIDS/cuDF, cuPyNumeric), HPC and Slurm workflows, distributed LLM training & inference (NeMo, Megatron), numerical optimization (cuOpt), and quantum computing (CUDA-Q) for accelerated-computing engineers.

Included Skills (89)

  1. NeMo Distributed Trainer — This tool assists developers in configuring and optimizing distributed training strategies within NeMo AutoModel to ensure efficient multi-GPU and multi-node model performance.
  2. Memory Optimization Assistant — This tool helps AI engineers optimize GPU memory usage by configuring selective or full activation recompute strategies within the Megatron Bridge framework.
  3. Parallelism Strategy Optimizer — This guide assists AI engineers in selecting and configuring optimal parallelism strategies for Megatron Bridge to maximize performance and hardware utilization.
  4. Megatron Resiliency Manager — This tool enables and manages fault tolerance, straggler detection, and automatic restart features to ensure stable training for Megatron Bridge users.
  5. GPU Portfolio Optimizer — This tool enables quantitative analysts to build, backtest, and optimize high-performance financial portfolios using NVIDIA-accelerated Mean-CVaR algorithms and GPU-based solvers.
  6. Megatron CPU Offloading — This skill helps developers optimize GPU memory usage in Megatron Bridge by configuring activation and optimizer state offloading to CPU memory.
  7. MoE Parallel Optimizer — This skill helps AI engineers optimize Megatron-Bridge performance by configuring expert-parallel communication overlap and dispatcher backends to reduce training latency.
  8. TAO Execution SDK — This SDK enables developers to manage, monitor, and scale NVIDIA TAO GPU training jobs across diverse cloud and local computing platforms.
  9. NeMo Job Launcher — This tool helps developers configure NeMo AutoModel job submissions for interactive environments, Slurm HPC clusters, and SkyPilot cloud-based execution platforms.
  10. TAO Docker Executor — This tool enables developers to execute TAO SDK jobs within local or remote Docker containers utilizing NVIDIA GPU acceleration for efficient model training.
  11. DALI Dynamic Assistant — This assistant helps developers write, review, and migrate deep learning data pipelines using the NVIDIA DALI imperative dynamic-mode API for GPU processing.
  12. Optimization Formulation Guide — This guide assists developers in translating complex business problems into structured LP, MILP, or QP mathematical models for efficient numerical optimization.
  13. Distributed Data Loader — This skill enables developers to efficiently load complex, sharded datasets into distributed cuPyNumeric ndarrays when standard built-in loaders are insufficient for custom layouts.
  14. CuTile Triton Converter — This tool automates the conversion and debugging of cuTile GPU kernels into Triton, assisting developers in porting high-performance code with optimized mapping.
  15. SLURM Cluster Executor — This tool enables AI agents to execute TAO training and inference jobs on remote SLURM GPU clusters using SSH, Pyxis, and Enroot containers.
  16. cuTile Kernel Assistant — This expert assistant helps developers write, debug, and optimize high-performance GPU kernels using the cuTile programming model for NVIDIA hardware.
  17. Brev GPU Orchestrator — This tool enables developers to manage NVIDIA TAO training and inference workloads by automating GPU instance deployment and job execution via the Brev CLI.
  18. Clinical ASR Builder — This tool assists clinical researchers in curating medical terminology, generating IPA tags, and synthesizing NeMo manifests for ASR model development.
  19. Clinical ASR Evaluator — This tool scores NeMo manifests to generate comprehensive KER leaderboards, assisting clinical AI researchers in evaluating and benchmarking speech recognition performance.
  20. Clinical ASR Bootstrapper — This tool automates the initial configuration and smoke testing of clinical ASR environments to help developers quickly bootstrap healthcare AI workflows.
  21. NeMo Model Onboarding — This guide assists developers in integrating new model architectures into NeMo AutoModel through structured implementation, registration, and validation workflows.
  22. NeMo Recipe Architect — This tool assists developers in creating, configuring, and validating NeMo AutoModel training and evaluation recipes through structured YAML and builder management.
  23. TileGym Kernel Integrator — This tool guides developers through the end-to-end process of registering, implementing, and benchmarking new cuTile GPU operators within the TileGym framework.
  24. HSB FPGA Flasher — This tool enables developers to safely flash FPGA firmware on HSB Lattice boards and Leopard Imaging VB1940 cameras connected to NVIDIA devkits.
  25. Holoscan Bridge Setup — This tool automates the end-to-end deployment, container configuration, and connectivity verification of the NVIDIA Holoscan Sensor Bridge for developers and engineers.
  26. Megatron Testing Framework — This tool assists developers in managing unit and functional tests, configuring recipe YAML files, and maintaining golden values for the Megatron-LM ecosystem.
  27. Sandbox Network Policy — Configures allowed egress endpoints for NemoClaw sandboxes to help developers manage network access rules and security policies for their applications.
  28. Lepton GPU Orchestrator — This tool enables developers to manage, dispatch, and monitor TAO container workloads on DGX Cloud Lepton GPU infrastructure without manual cluster configuration.
  29. CuTile Autotuning Assistant — This tool assists developers in designing, implementing, and optimizing CuTile autotuning kernels to achieve peak performance across diverse GPU architectures.
  30. Jetson Clock Customizer — This tool enables developers to configure Jetson CPU, GPU, and EMC clock behaviors by modifying pre-flash system files for optimized hardware performance.
  31. GPU Memory Optimizer — This skill provides techniques for reducing peak GPU memory usage and resolving OOM errors in Megatron Bridge for deep learning engineers.
  32. CUDA Graph Optimizer — This tool helps developers optimize training performance by configuring and validating CUDA graph implementations within Megatron Bridge for various model architectures.
  33. Megatron Bridge Recommender — This tool helps developers select and customize optimal Megatron Bridge training recipes based on specific model architectures, hardware configurations, and training objectives.
  34. GPU Dataframe Expert — Provides expert guidance for developers accelerating pandas workflows using NVIDIA cuDF and dask-cuDF to achieve high-performance GPU-based data processing and ETL.
  35. CuOpt Development Assistant — This tool assists developers in modifying, building, testing, and debugging the NVIDIA cuOpt codebase while strictly adhering to established project conventions.
  36. TAO Kubernetes Executor — This tool enables developers to deploy and manage NVIDIA TAO container jobs as scalable Kubernetes workloads with automated GPU resource scheduling.
  37. TAO AutoML Optimizer — Automates hyperparameter optimization for NVIDIA TAO models using advanced search algorithms to streamline training workflows and improve model performance for developers.
  38. VCN Gap Analyzer — This tool identifies weak classification samples for NVIDIA TAO VCN experiments to help engineers optimize decision thresholds and target data for augmentation.
  39. Nemotron Pipeline Orchestrator — This tool enables developers to plan, configure, and execute end-to-end Nemotron customization workflows, including training, optimization, and evaluation pipelines.
  40. Nemotron Safety Architect — This tool helps developers create custom safety policies, taxonomy, and inference prompts for NVIDIA Nemotron content-safety guardrails to ensure robust model governance.
  41. GPU Acceleration Expert — This skill helps developers transform CPU-bound Python code into high-performance GPU-accelerated applications using the NVIDIA RAPIDS ecosystem and CUDA-based libraries.
  42. Megatron-LM Slurm Orchestrator — This guide assists AI engineers in configuring and launching distributed Megatron-LM training jobs on SLURM clusters with optimized environment and network settings.
  43. Hierarchical Context Parallelism — This guide assists engineers in configuring and verifying hierarchical context parallelism within Megatron-Bridge to optimize large-scale model training performance.
  44. Megatron FSDP Optimizer — This guide assists engineers in configuring and verifying Megatron FSDP within Megatron-Bridge to optimize training performance and resolve memory issues.
  45. Dynamo Interconnect Validator — This tool validates RDMA and NVLink connectivity for Dynamo deployments to ensure reliable disaggregated serving performance for infrastructure engineers and operators.
  46. DeepStream Pipeline Architect — This skill assists developers in building and optimizing NVIDIA DeepStream 9.0 video analytics pipelines using Python and GStreamer-based inference integration.
  47. MoE Dispatcher Selector — This tool helps AI engineers select the optimal MoE token dispatcher based on hardware, EP degree, and specific model-family performance patterns.
  48. CuPyNumeric Migration Assessor — This tool evaluates NumPy codebases for cuPyNumeric compatibility, providing developers with actionable refactoring guidance and migration readiness verdicts for distributed GPU scaling.
  49. Jetson Inference Optimizer — This tool recommends optimal inference runtimes and memory configuration flags for LLM and VLM workloads running on NVIDIA Jetson hardware platforms.
  50. PhysicsNeMo Navigator — This tool helps researchers and developers navigate the PhysicsNeMo repository by identifying relevant models, datapipes, and examples for specific SciML and AI4Science tasks.
  51. Molecular Modeling Assistant — This tool provides a cloud-native Python interface for medicinal chemists to automate complex molecular simulations, drug discovery workflows, and protein-ligand modeling tasks.
  52. Holoscan Conda Installer — This tool automates the installation of Holoscan SDK v4.3+ into Conda environments for developers working on Linux x86_64 systems with CUDA 13.
  53. Nemotron Retrieval Assistant — This tool assists developers in planning, debugging, and deploying Nemotron embedding and reranking recipes by managing configurations, execution commands, and project workflows.
  54. CuOpt Routing Optimizer — This tool enables developers to solve complex vehicle routing problems efficiently using the NVIDIA cuOpt Python API for optimized logistics operations.
  55. Synthetic Dataset Generator — This tool helps data scientists and developers efficiently create custom synthetic datasets or build automated data generation pipelines for their specific projects.
  56. Autonomous RL Researcher — This agent automates iterative NeMo-RL experiment lifecycles, helping researchers conduct hypothesis testing, launch reproducible training runs, and track results within a git-based ledger.
  57. Session Memory Manager — This tool enables coding agents to persist and recover project context across disconnects, restarts, or handoffs, ensuring seamless continuity for complex, long-running development tasks.
  58. CuTile Kernel Converter — This tool automates the translation of Python cuTile GPU kernels into Julia cuTile.jl code, assisting developers in porting, debugging, and optimizing high-performance GPU applications.
  59. CuTile Kernel Optimizer — This tool helps developers systematically profile, diagnose, and iteratively tune cuTile GPU kernels to achieve peak performance within the TileGym framework.
  60. TAO Dataset Validator — This tool executes the tao-daft validate command to help developers verify the structure, schema, and integrity of NVIDIA TAO DAFT datasets.
  61. Jetson Health Monitor — This tool provides a comprehensive, read-only system snapshot to help developers and AI agents troubleshoot performance, thermal, and resource utilization on Jetson devices.
  62. Quantum Circuit Developer — This skill enables developers to design, simulate, and optimize NISQ quantum circuits specifically for Google Quantum AI hardware and various partner backends.
  63. Fluid Dynamics Simulator — This high-performance Python framework enables researchers and engineers to execute complex computational fluid dynamics simulations using advanced pseudospectral methods and parallel computing.
  64. Computational Resource Auditor — This tool assesses system hardware capabilities to provide strategic optimization recommendations for developers performing intensive scientific computing and large-scale data processing tasks.
  65. Serverless Cloud Orchestrator — This tool enables developers to deploy and scale Python applications, AI models, and GPU-accelerated workloads on a serverless cloud infrastructure.
  66. Polars Data Processor — This skill provides high-performance DataFrame manipulation and lazy evaluation capabilities to help data engineers and analysts optimize complex ETL and analytics pipelines.
  67. Quantum Circuit Framework — This framework enables developers to build, optimize, and execute quantum circuits across various hardware platforms, streamlining enterprise-grade quantum computing and algorithm development.
  68. Time Series Forecaster — This skill provides zero-shot univariate time series forecasting using Google's TimesFM model, enabling data analysts to generate accurate predictions without custom training.
  69. PEFT Fine-Tuning — Fine-tune LLMs efficiently using LoRA, QLoRA, and other PEFT methods, helping users adapt models to specific tasks on consumer GPUs.
  70. Vision Pipeline Specialist — Assists users with object detection, image segmentation, and deploying optimized computer vision pipelines using YOLO, SAM, and TensorRT.
  71. Azure Batch Automation — Automates Azure Batch job management for Java developers, simplifying large-scale parallel and HPC workload execution.
  72. TAO Dataset Converter — This tool automates the conversion of NVIDIA TAO DAFT datasets between supported formats to assist developers in managing their data pipelines efficiently.
  73. Quantum System Simulator — This tool provides comprehensive Python-based simulation capabilities for open and closed quantum systems, assisting physics researchers with dynamics, decoherence, and optical modeling.
  74. Modal Serverless GPU — Deploys and runs Python functions on cloud GPUs, enabling ML model deployment and inference without infrastructure management for developers.
  75. Triton Inference Server — Deploys AI models at scale, supporting multiple frameworks and hardware, benefiting data scientists and machine learning engineers.
  76. MoE Performance Optimizer — This workflow guides engineers through systematic MoE training optimization using the Three Walls framework to maximize throughput and resolve performance regressions.
  77. Megatron Training Orchestrator — This tool assists developers in executing and correlating Megatron-LM and Megatron Bridge training runs to ensure consistent loss curves and configuration parity.
  78. MoE Communication Optimizer — This tool helps performance engineers tune expert-parallel communication overlap in Megatron Bridge to maximize throughput for large-scale MoE model training.
  79. MoE Long-Context Optimizer — This guide provides performance optimization strategies for training long-context MoE models, helping engineers resolve memory constraints and maximize throughput in Megatron Bridge.
  80. Sequence Packing Optimizer — This skill assists developers in configuring and validating packed sequences and long-context training within Megatron-Bridge to optimize performance for LLM and VLM workloads.
  81. MoE VLM Trainer — This skill provides expert guidance for training MoE vision language models using Megatron Bridge, helping engineers optimize FSDP and 3D-parallel performance strategies.
  82. Communication Overlap Optimizer — This guide assists engineers in configuring and verifying TP, DP, and PP communication overlap settings within Megatron-Bridge to maximize training throughput.
  83. MoE Hardware Optimizer — This tool provides optimized training playbooks and throughput configurations for MoE models, assisting engineers in tuning performance across diverse hardware platforms.
  84. VCN Image Miner — This tool automates the DEFT embedding and mining workflow to help computer vision engineers augment training datasets with relevant nearest-neighbor source images.
  85. CuOpt Optimization CLI — This tool enables developers to solve linear, mixed-integer, and quadratic programming problems efficiently using the cuOpt command-line interface with MPS files.
  86. Nextflow Pipeline Engineer — This skill assists bioinformaticians and data scientists in building, debugging, and scaling reproducible data pipelines using Nextflow and the nf-core framework.
  87. PyTorch Lightning Assistant — This skill helps deep learning researchers and engineers organize PyTorch code, automate training workflows, and scale neural network models across multi-GPU or TPU environments.
  88. NemoClaw Sandbox Manager — Deploys and manages NemoClaw, NVIDIA's open-source sandbox, to securely run OpenClaw agents with policy-enforced network, filesystem, and inference controls.
  89. Pacsomatic Workflow Assistant — This toolkit assists bioinformaticians in validating inputs, generating samplesheets, and managing reproducible execution for the nf-core/pacsomatic tumor-normal analysis pipeline.