Amazon Redshift announces the general availability of automatic mounting of AWS Glue Data Catalog , making it easier for customers to run queries in their data lakes. You no longer have to create an external schema in Amazon Redshift to use the data lake tables cataloged in AWS Glue Data Catalog. Now, you can query data lake tables directly from Amazon Redshift Query Editor v2 or your favorite SQL editors.
Amazon EMR on EC2 announces support for native LDAP authentication
Amazon EMR is excited to announce a new feature that enables user authentication to EMR on EC2 clusters using Lightweight Directory Access Protocol (LDAP) based credentials. This feature allows administrators to configure EMR on EC2 clusters to authenticate corporate identities in their Active Directory (AD) using LDAP. With this launch, AD users are synced to the EMR on EC2 cluster automatically when LDAP authentication is enabled. This simplifies authentication to EMR clusters for administrators by eliminating the manual steps to sync users and/or implementing application-specific LDAP configuration.
Amazon SageMaker Canvas announces Document Queries powered by Amazon Textract
Amazon SageMaker Canvas now supports Document Queries, a ready-to-use model powered by Amazon Textract. Document Queries allows you to specify the data you want to extract from structured documents using natural language, without requiring prior knowledge of the document’s structure (table, form, fields, nested data). This eliminates the need for manual processing and searching within extracted data, saving you time and reducing human-error.
Amazon SageMaker Canvas supports sharing ML models with Amazon QuickSight
Amazon SageMaker Canvas now supports sharing machine learning (ML) models with Amazon QuickSight, enabling analysts to build models in Canvas and generate predictions to build dashboards in QuickSight. This extends the ML/Analytics integrated solution between Canvas and QuickSight for analysts to build models, generate predictions, enrich them with interactive dashboards, and use insights for effective business decisions, without writing a single line of code.
Amazon SageMaker Canvas expands data preparation with five new capabilities
Amazon SageMaker Canvas now supports five new data transforms, enabling you to better prepare and analyze your data before building machine learning (ML) models. Data is the foundation of machine learning and transforming raw data to make it suitable for ML model building and generating predictions is key to better insights. Starting today, SageMaker Canvas allows you to change the type of data in your columns between numeric, text, and datetime, while also displaying the associated feature for that data type such as binary and categorical. This gives you the flexibility to manually change the type of data in your columns based on the features. The ability to choose the right data type ensures data integrity and accuracy prior to building ML models. As an example, using a datetime data type ensures only valid dates are stored in that particular column.
Amazon SageMaker Canvas supports training ML models with different objective metrics
Amazon SageMaker Canvas now provides the ability to train machine learning (ML) models with different objective metrics, allowing you to gain a more comprehensive understanding on the model’s strengths and weaknesses. SageMaker Canvas is a visual interface that enables business analysts and citizen data scientists to generate accurate ML predictions on their own — without requiring any ML expertise or having to write a single line of code.
CloudWatch Application Insights adds monitoring for multi-app instance deployments
Amazon Web Services customers can now get detailed health metrics and analysis of their multi-application deployments residing in the same instance with Amazon CloudWatch Application Insights. CloudWatch Application Insights helps customers gain actionable insights for their application environment and AWS resources by making it easy to set up and monitor applications, recognize problems, and use data to make decisions.
Amazon SageMaker Canvas supports custom Amazon S3 output location for ML artifacts
Amazon SageMaker Canvas now supports the ability to provide a custom output location in Amazon S3 for machine learning (ML) artifacts, such as trained models, explainability reports, and prediction results allowing you to organize and structure your output directory in a way that aligns with your specific needs and preferences. SageMaker Canvas is a visual interface that enables business analysts and citizen data scientists to generate accurate ML predictions on their own — without requiring any ML expertise or having to write a single line of code.
AWS DataSync now supports copying data to and from Azure Blob Storage
AWS DataSync support for copying data to and from Azure Blob Storage is now generally available. Using DataSync, you can move your object data at scale between Azure Blob Storage and AWS Storage services such as Amazon S3. AWS DataSync supports writing to block blobs and can read from all blob types within Azure Blob Storage. It can also be used with Azure Data Lake Storage (ADLS) Gen 2.
Amazon EMR Serverless now supports storing logs in Amazon CloudWatch
Amazon EMR Serverless is a serverless option that makes it simple for data analysts and engineers to run open-source big data analytics frameworks like Apache Spark and Apache Hive without configuring, managing, and scaling clusters or servers. Starting today, you can store logs for your EMR Serverless Spark and Hive applications in Amazon CloudWatch.