Amazon SageMaker Autopilot, a low-code machine learning (ML) service which automatically builds, trains, and tunes the best ML models based on your data, is now integrated with Amazon SageMaker Pipelines, the first purpose-built continuous integration and continuous delivery (CI/CD) service for ML. This enables the automation of an end-to-end flow of building ML models using SageMaker Autopilot and integrating models into subsequent CI/CD steps.
Introducing Amazon SageMaker support for shadow testing
Amazon SageMaker supports shadow testing to help you validate performance of new machine learning (ML) models by comparing them to production models. With shadow testing, you can spot potential configuration errors and performance issues before they impact end users. SageMaker eliminates weeks of time spent building infrastructure for shadow testing, so you can release models to production faster.
Amazon SageMaker Studio now supports real time collaboration
Amazon SageMaker Studio is a fully integrated development environment (IDE) for machine learning (ML) that enables ML practitioners to perform every step of the machine learning workflow, from preparing data to building, training, tuning, and deploying models. Today, we are announcing new capabilities in SageMaker Studio to accelerate real time collaboration across ML teams.
Amazon SageMaker Studio now supports automatic conversion of notebook code to production-ready jobs
Amazon SageMaker Studio is a fully integrated development environment (IDE) for machine learning (ML) that enables ML practitioners to perform every step of the machine learning workflow, from preparing data to building, training, tuning, and deploying models. Today, we’re excited to announce a new capability in SageMaker Studio notebooks that enables automatic conversion of notebook code to production-ready jobs.
Amazon SageMaker Data Wrangler now provides built-in data preparation in notebooks
Amazon SageMaker Data Wrangler reduces the time it takes to aggregate and prepare data for ML from weeks to minutes With Data Wrangler, you can simplify the process of data preparation and feature engineering, and complete each step of the data preparation workflow, including data selection, visualization, cleansing, and preparation from a low-code visual interface. Many ML practitioners want to explore datasets directly in notebooks to spot potential data-quality issues, like missing information, extreme values, skewed datasets, or biases, so they can correct those issues to prepare data for training ML model faster. ML practitioners can spend weeks writing boilerplate code to visualize and examine different parts of their dataset to identify and fix potential issues.
Amazon GuardDuty RDS Protection now in preview
Amazon GuardDuty now offers threat detection for Amazon Aurora to identify potential threats to data stored in Aurora databases. Amazon GuardDuty RDS Protection profiles and monitors access activity to existing and new databases in your account, and uses tailored machine learning models to accurately detect suspicious logins to Aurora databases. Once a potential threat is detected, GuardDuty generates a security finding that includes database details and rich contextual information on the suspicious activity, is integrated with Aurora for direct access to database events without requiring you to modify your databases, and is designed to not affect database performance.
Amazon S3 Access Points can now be used to securely delegate access permissions for shared datasets to other AWS accounts
Amazon S3 Access Points simplify data access for any AWS service or customer application that stores data in S3 buckets. With S3 Access Points, you create unique access control policies for each access point to more easily control access to shared datasets. Now, bucket owners are able to authorize access via access points created in other accounts. In doing so, bucket owners always retain ultimate control over data access, but can delegate responsibility for more specific IAM-based access control decisions to the access point owner. This allows you to securely and easily share datasets with thousands of applications and users, and at no additional cost.
Introducing the Amazon EC2 Spot Ready Software Products
The new Amazon EC2 Spot Ready specialization helps customers identify validated AWS Partner software products that support Amazon EC2 Spot Instances, a compute purchase option that allows customers to utilize spare EC2 capacity at a discounted price from on demand (up to 90%). Amazon EC2 Spot Ready ensures that customers have a well-architected and cost-optimized solution to help them benefit from EC2 Spot savings for their workloads.
Introducing Amazon Managed Streaming for Apache Kafka (MSK) Delivery Partners
Amazon Web Services (AWS) is incredibly excited to announce the new Amazon MSK Service Delivery specialization for AWS partners that help customers migrate and build real-time streaming analytics solutions with fully managed Apache Kafka. Amazon MSK provisions your servers, configures your Apache Kafka clusters, replaces servers when they fail, orchestrates server patches and upgrades, architects clusters for high availability, ensures data is durably stored and secured, sets up monitoring and alarms, and runs scaling to support load changes. With MSK Serverless, getting started with Apache Kafka is even easier. It automatically provisions and scales compute and storage resources and offers throughput-based pricing, so you can use Apache Kafka on demand and pay for the data you stream and retain.
Introducing AWS Glue Delivery
We are excited to announce the new AWS Glue Delivery specialization, which validates AWS Partners with deep expertise and proven success delivering AWS Glue for data integration, data pipeline, and data catalogue use cases. AWS Glue is a scalable, serverless data integration service that makes it easy to discover, prepare, and combine data for analytics, machine learning, and application development. With the ability to scale on demand, AWS Glue helps customers focus on high-value activities that maximize the value of their data.