AWS Glue announces the preview of AWS Glue Data Quality, a new capability that automatically measures and monitors data lake and data pipeline quality. AWS Glue is a serverless, scalable data integration service that makes it more efficient to discover, prepare, move, and integrate data from multiple sources. Managing data quality is manual and time-consuming. You must set up data quality rules and validate your data against these rules on a recurring basis, also writing code to set up alerts when quality deteriorates. Analysts must manually analyze data, write rules, and then write code to implement these rules.
Amazon Redshift data sharing now supports centralized access control with AWS Lake formation (Preview)
Amazon Redshift data sharing enables you to efficiently share live data across Amazon Redshift data warehouses. Amazon Redshift now supports simplified governance of Amazon Redshift data sharing by enabling you to use AWS Lake Formation to centrally manage permissions on data being shared across your organization. With the new Amazon Redshift data sharing managed by AWS Lake Formation, you can view, modify, and audit permissions on the tables and views in the Redshift datashares using Lake Formation APIs and the AWS Console, and allow the Redshift datashares to be discovered and consumed by other Redshift data warehouses.
Introducing new ML governance tools for Amazon SageMaker
Today, we are excited to announce three new purpose-built tools for Amazon SageMaker to improve governance of your machine learning (ML) projects with simplified access control and enhanced transparency across your ML model’s lifecycle. With Amazon SageMaker Role Manager, you can define minimum permissions for users in minutes and onboard new users faster. SageMaker Role Manager simplifies the permission setting for ML activities and automatically generates a custom policy based on your specific needs.
Deploy SageMaker Data Wrangler for real-time and batch inference and additional configurations to processing jobs
Today, we are excited to announce support for deploying data preparation flows created in Data Wrangler to real-time and batch serial inference pipelines, and additional configurations for Data Wrangler processing jobs in Amazon SageMaker Data Wrangler.
Introducing AWS AI Service Cards – a new resource for responsible AI
We are excited to announce AWS AI Service Cards, a new resource to increase transparency and help customers better understand our AWS AI services, including how to use them in a responsible way. AI service cards are a form of responsible AI documentation that provides customers with a single place to find information on the intended use cases and limitations, responsible AI design choices, and best deployment and operation practices for our AI Services. They are part of a comprehensive development process we undertake to build our services in a responsible way with fairness and bias, robustness, explainability, governance, transparency, privacy, and security in mind.
Amazon SageMaker Data Wrangler now supports over 40 third-party applications as data sources
Today, AWS announces the general availability of Amazon SageMaker Data Wrangler support for over 40 third party applications as data sources for machine learning (ML) through the integration with Amazon AppFlow. Amazon SageMaker Data Wrangler reduces the time it takes to aggregate and prepare data for machine learning (ML) from weeks to minutes. Preparing high quality data for ML is often complex and time consuming as it requires aggregating data across various sources and formats using different tools. With SageMaker Data Wrangler, you can explore and import data from a variety of popular sources, such as Amazon S3, Amazon Athena, Amazon Redshift, Snowflake, Databricks and Salesforce Customer Data Platform. Starting today, we are making it easier for customers to aggregate data for ML from over 40 third-party application data sources, including Salesforce Marketing, SAP, Google Analytics, LinkedIn and more via Amazon AppFlow.
Amazon AppFlow now supports over 50 Connectors
Amazon AppFlow announces the release of 22 new data connectors. With this launch, Amazon AppFlow now supports data connectivity to over 50 applications. Amazon AppFlow is a fully managed integration service that enables you to securely transfer data between Software-as-a-Service (SaaS) applications and AWS services like Amazon S3 and Amazon Redshift. As enterprises increasingly rely on SaaS services for mission-critical workflows, they face the challenge of collecting data from a growing ecosystem of services into a centralized location to derive business insights using analytics and machine learning. With Amazon AppFlow, you can easily set up data flows in minutes without writing code.
AWS Machine Learning University announces educator enablement program for higher education
AWS Machine Learning University is now providing a free educator enablement program that prioritizes U.S. community colleges, Minority Serving Institutions (MSIs), and Historically Black Colleges and Universities (HBCUs). Educators can leverage these tools to launch stand-alone courses, certificates, or full degrees in data management (DM), artificial intelligence (AI), and machine learning (ML). The goal is to make early-career DM/AI/ML jobs more accessible to a broader and more diverse student population. The program offers a suite of ready-to-use tools to faculty, including a library of ready-to-teach DM/AI/ML educational materials, free computing capacity, and comprehensive faculty professional development built around MLU, Amazon’s own internal training program for ML practitioners.
Amazon Redshift now supports auto-copy from Amazon S3
Amazon Redshift launches the preview of auto-copy support to simplify data loading from Amazon S3 into Amazon Redshift. You can now setup continuous file ingestion rules to track your Amazon S3 paths and automatically load new files without the need for additional tools or custom solutions.
Amazon SageMaker JumpStart now enables you to more easily share ML artifacts within your organization
Amazon SageMaker JumpStart now enables you to more easily share machine learning (ML) artifacts, including notebooks and models, across your organization to accelerate model building and deployment. Amazon SageMaker JumpStart is an ML hub that accelerates your ML journey with built-in algorithms and pretrained models from popular model hubs, such as Hugging Face, and end-to-end solutions that solve common use cases.