Data Science & Analytics Bundle | Prompeteer.ai

Data analysis, SQL mastery, visualization, statistical modeling, and data pipeline skills for data professionals.

Included Skills (90)

  1. Data Engineering Expert — This skill helps data scientists and analysts by building scalable data pipelines, ETL/ELT systems, and robust data infrastructure.
  2. Jupyter Notebook Assistant — Assists data scientists and researchers with interactive data analysis, visualization, and reproducible research using Jupyter notebooks.
  3. Agentic Analytics Architect — This skill helps data architects design and deploy secure, governed analytics solutions for distributed data across multi-cloud and on-premises environments.
  4. Embedded Analytics Database — This skill helps data professionals use DuckDB, an embedded analytical database, for local data exploration and ETL pipelines.
  5. BigQuery BigFrames Assistant — This tool generates optimized Python code using the BigFrames API to help data scientists perform scalable, pandas-style analysis directly within BigQuery.
  6. BigQuery AI Analyst — This skill empowers data analysts to perform advanced machine learning, predictive forecasting, and generative AI tasks directly within BigQuery using SQL.
  7. BigQuery Data Analyzer — Analyze large datasets using Google BigQuery, enabling data professionals to run SQL queries and build machine learning models.
  8. Portable Analytics Assistant — Provides expert guidance for Ibis, enabling developers to write portable analytics code that runs on various SQL backends.
  9. Snowflake Development Assistant — Assists data engineers and analysts in developing Snowflake SQL, data pipelines, and Snowpark Python applications efficiently and effectively.
  10. SQL Query Generator — This tool transforms natural language requirements into optimized SQL queries, helping analysts and engineers efficiently retrieve data across various database platforms.
  11. Distributed Computing Engine — This tool enables data scientists and engineers to scale pandas and NumPy workflows beyond memory limits across single machines or distributed clusters.
  12. Evidence Dashboard Assistant — Provides expert guidance for Evidence, helping developers build data reports as code and create self-service analytics dashboards from SQL and Markdown.
  13. clickhouse — The skill executes real‑time analytical queries on massive datasets through a familiar SQL interface, offering advanced aggregation functions and high‑throughput inserts. It supports materialized views for pre‑computed rollups, enabling efficient event tracking, time‑series analytics, and ad‑hoc exploration of billions of rows. Data analysts, engineers, and product teams benefit from fast, scalable analytics on large volumes of data.
  14. MotherDuck SQL Assistant — Provides expert guidance for MotherDuck, helping developers perform SQL analytics on cloud data and build hybrid local-cloud data pipelines.
  15. Spark Data Processing — Process large-scale data using Apache Spark for data analysis, ETL pipelines, and distributed computations, assisting data engineers and analysts.
  16. Product Analytics — Use when defining product KPIs, building metric dashboards, running cohort or retention analysis, or interpreting feature adoption trends across product stages.
  17. Metabase Analytics Tool — Metabase helps users create dashboards and visualizations from data, enabling data-driven decision-making within organizations, even without SQL knowledge.
  18. GPU Dataframe Expert — Provides expert guidance for developers accelerating pandas workflows using NVIDIA cuDF and dask-cuDF to achieve high-performance GPU-based data processing and ETL.
  19. Statistical Visualization Assistant — This skill enables data scientists to generate publication-quality statistical graphics and complex multi-panel figures by leveraging Seaborn's seamless integration with pandas data structures.
  20. Scientific Visualization Assistant — This tool helps researchers create, audit, and format publication-ready scientific figures while ensuring data integrity, accessibility, and compliance with specific journal submission requirements.
  21. Data Migration — The Data Migration skill analyzes source and target schemas, maps fields, and processes data in batches while validating outcomes and planning cutover to ensure accurate transfers. It transforms schemas for seamless integration between databases, providing reliable, comprehensive migration support. Database administrators and data engineers rely on it to move data efficiently and accurately.
  22. Public Database Retriever — This tool enables developers and researchers to perform reproducible, auditable data lookups across authoritative public APIs with precise filtering and full provenance tracking.
  23. Single-Cell Analysis Toolkit — This toolkit provides a comprehensive Python pipeline for single-cell RNA-seq analysis, assisting researchers with quality control, clustering, and visualization of complex biological datasets.
  24. Big Data Analyst — This skill empowers data professionals to perform high-speed analysis, visualization, and machine learning on massive tabular datasets that exceed available system memory.
  25. Product Analytics Advisor — Provides expert product analytics guidance, enabling product teams to define metrics, analyze funnels, and make data-driven decisions.
  26. Customer Data Routing — This skill helps developers and marketers track and route customer data to various analytics and marketing tools using Segment's CDP.
  27. ECharts Visualization Generator — Create interactive data visualizations using Apache ECharts for developers building charts and dashboards in JavaScript applications.
  28. Pandas Data Assistant — This skill assists users with loading, cleaning, transforming, and analyzing tabular data using the pandas library in Python.
  29. MySQL Database Management — This skill helps developers manage MySQL databases, covering installation, SQL queries, schema design, and client integration with Node.js and Python.
  30. Pipeline Health Analyzer — This tool evaluates sales opportunity quality and win probability to help sales managers prioritize pipeline reviews and improve data hygiene.
  31. Polars Data Processor — This skill provides high-performance DataFrame manipulation and lazy query optimization to help data engineers and analysts build efficient, scalable Python data pipelines.
  32. Video Reasoning Annotator — This pipeline transforms raw video content into structured Chain-of-Thought training datasets, helping researchers and developers generate high-quality, annotated video understanding data.
  33. TimesFM Time Series — Forecast time series data using Google's TimesFM foundation model for zero-shot prediction, assisting data scientists and analysts.
  34. Flink Stream Processor — This skill processes real-time data streams using Apache Flink, assisting users with real-time analytics and event stream processing.
  35. A/B Test Designer — This skill assists users in planning, designing, and implementing effective A/B tests and experiments for data-driven decision-making.
  36. RAG Pipeline Architect — This tool helps developers design, optimize, and evaluate production-grade RAG pipelines by utilizing data-driven chunking strategies and rigorous retrieval performance metrics.
  37. Neo4j Graph Database — Assists developers in utilizing Neo4j, a graph database, for managing connected data and building graph-based applications.
  38. D3 Visualization Assistant — This skill helps developers build custom interactive data visualizations using D3.js for charts, graphs, maps, and complex diagrams.
  39. Advanced Data Scientist — This skill provides expert statistical modeling, causal inference, and machine learning workflows to help data scientists build robust, production-grade predictive systems.
  40. Dagster Pipeline Orchestration — This skill helps data engineers and scientists define, manage, and orchestrate data pipelines using Dagster's software-defined asset framework.
  41. Analytics Quality Framework — This skill helps data engineers and analysts establish robust testing, automated monitoring, and incident response protocols for analytics models and dashboards.
  42. Cohort Performance Analyzer — This tool enables business analysts to evaluate performance trends and diagnose conversion patterns across various customer segments, booking vintages, and acquisition channels.
  43. Analytics Instrumentation Architect — This skill helps product and data teams design, govern, and maintain standardized event tracking plans for robust analytics pipelines.
  44. Synthetic Dataset Generator — This tool enables data scientists and developers to efficiently create custom synthetic datasets and build automated data generation pipelines for their projects.
  45. Dimensionality Reduction Tool — This tool enables data scientists to perform efficient nonlinear dimensionality reduction and generate high-quality embeddings for visualization and advanced machine learning workflows.
  46. Web Data Extractor — This skill extracts structured data from web pages, handling static HTML and JavaScript-rendered content, to help users automate data collection.
  47. Gene Network Inference — Arboreto enables bioinformaticians to infer gene regulatory networks from transcriptomics data by identifying transcription factor-target gene relationships using scalable ensemble regression algorithms.
  48. Matplotlib Visualization Expert — This skill provides comprehensive guidance on creating highly customized, publication-quality static and interactive plots for data scientists and researchers using Python.
  49. Malloy Language Expert — Provides expert guidance for Malloy, Google's data language, helping developers write models, build queries, and explore data effectively.
  50. Cloud Agentic Architect — This tool helps data scientists and engineers design and implement robust, multi-product agentic architectures on Google Cloud using industry best practices.
  51. Mapbox Application Builder — Build interactive map-based applications with Mapbox GL JS, assisting developers with geocoding, navigation, and geospatial data visualization.
  52. Superset Data Explorer — Guides users in deploying, connecting, visualizing, and accessing data with Apache Superset, an open-source data exploration platform.
  53. PostHog Analytics Platform — This skill helps product managers and developers implement and utilize PostHog for product analytics, feature flags, and session replay.
  54. Mixpanel Analytics Integration — This skill helps developers integrate Mixpanel analytics into web and mobile applications for tracking user behavior and product performance.
  55. Loyalty Analytics Engine — This tool empowers marketing and operations teams to analyze member behavior, segment audiences, and evaluate experiment performance through data-driven insights.
  56. Cancer Imaging Explorer — This tool enables researchers to query and download large-scale public cancer imaging datasets from the NCI Imaging Data Commons for AI development.
  57. Data Enrichment Optimizer — This skill helps RevOps and data engineering teams maximize data quality and minimize costs by optimizing provider selection and waterfall enrichment sequences.
  58. Data Lakehouse Architect — This tool assists architects in designing secure, governed, and borderless open data lakehouse architectures that seamlessly integrate advanced agentic AI capabilities.
  59. Kysely Query Builder — Write type-safe SQL queries using the Kysely query builder, assisting developers who need TypeScript safety and SQL power.
  60. Spanner Database Expert — Provides developers and administrators with guidance on schema design, performance optimization, and secure management of Google Cloud Spanner database instances and resources.
  61. Product Metrics Architect — This tool helps product managers and data analysts design actionable metrics dashboards, define key performance indicators, and establish effective data monitoring strategies.
  62. Scikit-learn Assistant — Helps data scientists and machine learning engineers build, evaluate, and deploy models using the scikit-learn library.
  63. SQLite Database Expert — This skill helps developers learn SQLite CLI, Node.js, Python integration, and best practices for database management.
  64. Earth2Studio Datasource Creator — This tool assists developers in implementing, testing, and validating new Earth2Studio data source wrappers to integrate remote data stores into the platform.
  65. Deterministic Forecast Builder — This tool assists researchers in constructing single-member weather forecast inference scripts using the Earth2Studio framework for efficient climate modeling and analysis.
  66. Statistical Research Assistant — This skill provides researchers with comprehensive statistical analysis, including test selection, assumption verification, power calculations, and professional APA-style reporting for experimental data.
  67. Time Series Analyst — This skill provides advanced machine learning algorithms for time series analysis, helping data scientists perform classification, forecasting, and anomaly detection on sequential data.
  68. Vision Pipeline Specialist — Assists users with object detection, image segmentation, and deploying optimized computer vision pipelines using YOLO, SAM, and TensorRT.
  69. Data Lineage Summarizer — This tool generates intuitive Markdown reports of Google Cloud data lineage graphs to help data engineers debug quality issues and trace complex asset provenance.
  70. Access Database Manager — This skill helps users build, manage, and migrate Microsoft Access databases, queries, forms, reports, and VBA automation.
  71. Bioinformatics Analysis Toolkit — This toolkit provides comprehensive Python tools for sequence manipulation, phylogenetic analysis, and microbial ecology statistics to assist researchers in processing complex biological datasets.
  72. Analytics Tracking Assistant — Assists users in setting up, improving, and auditing analytics tracking and measurement for actionable marketing and product insights.
  73. Startup Data Strategist — Provides strategic guidance for startup founders and CDOs on AI data rights, architectural decisions, asset valuation, and organizational scaling for data-driven growth.
  74. Streamlit App Assistant — Provides expert assistance for Streamlit, enabling developers to build interactive data applications and dashboards using Python, without requiring frontend expertise.
  75. Geospatial Analysis Engine — This comprehensive toolkit empowers researchers and data scientists to perform advanced remote sensing, GIS operations, and spatial machine learning across diverse Earth observation domains.
  76. Apache Arrow Guide — Provides expert assistance for Apache Arrow, enabling developers to utilize its high-performance columnar format for data interchange and processing.
  77. LLM Observability Proxy — Helicone logs LLM requests, enables caching/rate-limiting, and provides cost analytics, benefiting developers using OpenAI, Anthropic, and other LLM providers.
  78. WandB Experiment Tracker — Tracks machine learning experiments, optimizes hyperparameters, and manages artifacts for data scientists and machine learning engineers.
  79. Data Version Control — DVC tracks large datasets and ML models, builds reproducible pipelines, and enables experiment tracking for data scientists and ML engineers.
  80. Startup Data Strategist — Provides strategic guidance for startup founders and CDOs on AI data rights, architectural selection, asset valuation, and organizational scaling for data-driven growth.
  81. Alembic Migration Manager — Manages database schema migrations using Alembic, enabling developers to version control and automate database changes efficiently.
  82. Cohort Retention Analyst — This tool analyzes user engagement and retention patterns to help product teams identify churn trends, feature adoption rates, and long-term user behavior insights.
  83. Autonomous Artifact Optimizer — This tool autonomously improves code, prompts, and data pipelines through iterative hypothesis tree refinement to achieve superior performance without overfitting to evaluation sets.
  84. Bioinformatics Query Tool — This tool provides researchers with unified command-line and Python access to over twenty genomic databases for rapid sequence analysis and data retrieval.
  85. Excel Data Processor — This skill processes Excel and CSV files, enabling data analysis, transformation, and report generation for users working with tabular data.
  86. Data Strategy Auditor — This tool provides a rigorous, decision-driven framework for leaders to pressure-test data architecture, training initiatives, and productization plans before committing resources.
  87. GPU Acceleration Optimizer — This skill accelerates scientific Python workloads on NVIDIA hardware by implementing and benchmarking GPU-optimized libraries to ensure measurable performance gains for data-intensive applications.
  88. Prefect Workflow Orchestration — This skill helps data scientists and engineers orchestrate Python data pipelines with Prefect, enabling scheduling, monitoring, and retries.
  89. Pinecone Vector Database — This skill helps AI developers use Pinecone, a managed vector database, to build semantic search and RAG applications.
  90. Airflow Workflow Orchestration — This skill helps data engineers and scientists automate, schedule, and monitor their data pipelines using Apache Airflow.