お手軽DEA-C01問題集PDFのベスト問題集を使おう!高得点目指すならここ [Q17-Q35]

Share

お手軽DEA-C01問題集PDFのベスト問題集を使おう!高得点目指すならここ

SnowPro Advanced DEA-C01試験と認定テストエンジン


Snowflake DEA-C01 認定試験の出題範囲:

トピック出題範囲
トピック 1
  • ストレージとデータ保護: このトピックでは、データ回復機能の実装と、Snowflake のタイム トラベルおよびマイクロ パーティションの理解をテストします。エンジニアは、クローン作成を通じて新しい環境を作成し、データ保護を確保する能力について評価され、Snowflake データの整合性とアクセス性を維持するための必須スキルが強調されます。
トピック 2
  • データ変換: SnowPro Advanced: Data Engineer 試験では、ユーザー定義関数 (UDF)、外部関数、およびストアド プロシージャの使用スキルを評価します。半構造化データを処理し、変換に Snowpark を活用する能力を評価します。このセクションでは、Snowflake エンジニアがデータ操作タスクに不可欠な Snowflake 環境内でデータを効果的に変換できることを保証します。
トピック 3
  • パフォーマンスの最適化: このトピックでは、Snowflake でパフォーマンスの低いクエリを最適化およびトラブルシューティングする能力を評価します。候補者は、最適なソリューションの構成、キャッシュの利用、およびデータ パイプラインの監視に関する知識を証明する必要があります。このトピックでは、エンジニアが特定のシナリオに基づいてパフォーマンスを向上できるようにすることに重点を置いています。これは、Snowflake データ エンジニアとソフトウェア エンジニアにとって重要です。
トピック 4
  • セキュリティ: DEA-C01 テストのセキュリティ トピックでは、システム ロールの管理やデータ ガバナンスなど、Snowflake セキュリティの原則について学習します。Snowflake データ エンジニアとソフトウェア エンジニアにとって安全なデータ環境を維持するために重要な、データを保護し、ポリシーに準拠する能力を測定します。
トピック 5
  • データ移動: Snowflake データ エンジニアとソフトウェア エンジニアは、Snowflake でのデータの読み込み、取り込み、トラブルシューティングの能力に基づいて評価されます。継続的なデータ パイプラインの構築、コネクタの構成、データ共有ソリューションの設計に関するスキルを評価します。

 

質問 # 17
A data engineering team is using an Amazon Redshift data warehouse for operational reporting.
The team wants to prevent performance issues that might result from long- running queries. A data engineer must choose a system table in Amazon Redshift to record anomalies when a query optimizer identifies conditions that might indicate performance issues.
Which table views should the data engineer use to meet this requirement?

  • A. STL_QUERY_METRICS
  • B. STL_ALERT_EVENT_LOG
  • C. STL_USAGE_CONTROL
  • D. STL_PLAN_INFO

正解:B

解説:
https://docs.aws.amazon.com/redshift/latest/dg/cm_chap_system-tables.html STL_ALERT_EVENT_LOG table view to meet this requirement. This system table in Amazon Redshift is designed to record anomalies when a query optimizer identifies conditions that might indicate performance issues.


質問 # 18
Snowflake supports using key pair authentication for enhanced authentication security as an alterna-tive to basic authentication (i.e. username and password). Select the list of SnowFlake Clients sup-port the same?
[Select All that Apply]

  • A. Node.js
  • B. SnowFlake Connector for Spark
  • C. SnowSQL
  • D. SnowCD
  • E. Go Driver

正解:A、B、C、E


質問 # 19
A data engineer has a one-time task to read data from objects that are in Apache Parquet format in an Amazon S3 bucket. The data engineer needs to query only one column of the data.
Which solution will meet these requirements with the LEAST operational overhead?

  • A. Prepare an AWS Glue DataBrew project to consume the S3 objects and to query the required column.
  • B. Use S3 Select to write a SQL SELECT statement to retrieve the required column from the S3 objects.
  • C. Configure an AWS Lambda function to load data from the S3 bucket into a pandas dataframe.
    Write a SQL SELECT statement on the dataframe to query the required column.
  • D. Run an AWS Glue crawler on the S3 objects. Use a SQL SELECT statement in Amazon Athena to query the required column.

正解:B

解説:
https://docs.aws.amazon.com/AmazonS3/latest/userguide/storage-inventory-athena-query.html S3 Select allows you to retrieve a subset of data from an object stored in S3 using simple SQL expressions. It is capable of working directly with objects in Parquet format.


質問 # 20
A retail company uses Amazon Aurora PostgreSQL to process and store live transactional data.
The company uses an Amazon Redshift cluster for a data warehouse.
An extract, transform, and load (ETL) job runs every morning to update the Redshift cluster with new data from the PostgreSQL database. The company has grown rapidly and needs to cost optimize the Redshift cluster.
A data engineer needs to create a solution to archive historical data. The data engineer must be able to run analytics queries that effectively combine data from live transactional data in PostgreSQL, current data in Redshift, and archived historical data. The solution must keep only the most recent 15 months of data in Amazon Redshift to reduce costs.
Which combination of steps will meet these requirements? (Choose two.)

  • A. Configure the Amazon Redshift Federated Query feature to query live transactional data that is in the PostgreSQL database.
  • B. Create a materialized view in Amazon Redshift that combines live, current, and historical data from different sources.
  • C. Schedule a monthly job to copy data that is older than 15 months to Amazon S3 Glacier Flexible Retrieval by using the UNLOAD command. Delete the old data from the Redshift cluster. Configure Redshift Spectrum to access historical data from S3 Glacier Flexible Retrieval.
  • D. Schedule a monthly job to copy data that is older than 15 months to Amazon S3 by using the UNLOAD command. Delete the old data from the Redshift cluster. Configure Amazon Redshift Spectrum to access historical data in Amazon S3.
  • E. Configure Amazon Redshift Spectrum to query live transactional data that is in the PostgreSQL database.

正解:A、D

解説:
Choice A ensures that live transactional data from PostgreSQL can be accessed directly within Redshift queries.
Choice C archives historical data in Amazon S3, reducing storage costs in Redshift while still making the data accessible via Redshift Spectrum.


質問 # 21
When would a Data engineer use table with the flatten function instead of the lateral flatten combination?

  • A. When table withFLATTENis acting like a sub-query executed for each returned row
  • B. WhenTABLE with FLATTENrequires no additional source m the from clause to refer to
  • C. When TABLE with FLATTENrequires another source in the from clause to refer to
  • D. Whenthe LATERALFLATTENcombination requires no other source m the from clause to refer to

正解:C

解説:
Explanation
The TABLE function with the FLATTEN function is used to flatten semi-structured data, such as JSON or XML, into a relational format. The TABLE function returns a table expression that can be used in the FROM clause of a query. The TABLE function with the FLATTEN function requires another source in the FROM clause to refer to, such as a table, view, or subquery that contains the semi-structured data. For example:
SELECT t.value:city::string AS city, f.value AS population FROM cities t, TABLE(FLATTEN(input => t.value:population)) f; In this example, the TABLE function with the FLATTEN function refers to the cities table in the FROM clause, which contains JSON data in a variant column named value. The FLATTEN function flattens the population array within each JSON object and returns a table expression with two columns: key and value.
The query then selects the city and population values from the table expression.


質問 # 22
Which ones are the false statements about Materialized Views?

  • A. Snowflake does not allow users to truncate materialized views.
  • B. A materialized view can also be used as the data source for a subquery.
  • C. Snowflake does not allow standard DML (e.g. INSERT, UPDATE, DELETE) on ma-terialized views.
  • D. Materialized views are first-class account objects.
  • E. Clustering a subset of the materialized views on a table tends to be more cost-effective than clustering the table itself.
  • F. Materialized views can be secure views.

正解:D

解説:
Explanation
Materialized views are first-class Database objects & rest of the understandings are true.


質問 # 23
Which methods can be used to create a DataFrame object in Snowpark? (Select THREE)

  • A. session.read.json{)
  • B. session,table()
  • C. DataFraas.writeO
  • D. session.sql()
  • E. session.builder()
  • F. session.jdbc_connection()

正解:A、B、D

解説:
Explanation
The methods that can be used to create a DataFrame object in Snowpark are session.read.json(), session.table(), and session.sql(). These methods can create a DataFrame from different sources, such as JSON files, Snowflake tables, or SQL queries. The other options are not methods that can create a DataFrame object in Snowpark. Option A, session.jdbc_connection(), is a method that can create a JDBC connection object to connect to a database. Option D, DataFrame.write(), is a method that can write a DataFrame to a destination, such as a file or a table. Option E, session.builder(), is a method that can create a SessionBuilder object to configure and build a Snowpark session.


質問 # 24
A company uses Amazon S3 to store semi-structured data in a transactional data lake. Some of the data files are small, but other data files are tens of terabytes.
A data engineer must perform a change data capture (CDC) operation to identify changed data from the data source. The data source sends a full snapshot as a JSON file every day and ingests the changed data into the data lake.
Which solution will capture the changed data MOST cost-effectively?

  • A. Use an open source data lake format to merge the data source with the S3 data lake to insert the new data and update the existing data.
  • B. Create an AWS Lambda function to identify the changes between the previous data and the current data. Configure the Lambda function to ingest the changes into the data lake.
  • C. Ingest the data into an Amazon Aurora MySQL DB instance that runs Aurora Serverless. Use AWS Database Migration Service (AWS DMS) to write the changed data to the data lake.
  • D. Ingest the data into Amazon RDS for MySQL. Use AWS Database Migration Service (AWS DMS) to write the changed data to the data lake.

正解:A


質問 # 25
Dominic, a Data Engineer wants to resume the pipe named stalepipe3 which got stale after 14 days. To do the same, he called the SYSTEM$PIPE_FORCE_RESUME function select sys-tem$pipe_force_resume('snowmydb.mysnowschema.stalepipe3','staleness_check_override'); Let's say If the pipe is resumed 16 days after it was paused, what will happened to the event notifi-cation that were received on the first and second days after the pipe was paused?

  • A. All the events get processed from day 1 if the PURGE properties in the PIPE object definition set to be FALSE initially.
  • B. Pipe maintains Metadata history of files for 64 days, so in this scenarios Snowpipe pro-cessed all the event notifications that were received for 16 days or so.
  • C. Once the Pipe got stale, all the events got purged automatically & pipe needs to be rec-reated with modified properties.
  • D. Snowpipe generally skips any event notifications that were received on the first and second days after the pipe was paused.

正解:D

解説:
Explanation
When a pipe is paused, event messages received for the pipe enter a limited retention period. The period is 14 days by default. If a pipe is paused for longer than 14 days, it is considered stale.
To resume a stale pipe, a qualified role must call the SYSTEM$PIPE_FORCE_RESUME function and input the STALENESS_CHECK_OVERRIDE argument. This argument indicates an under-standing that the role is resuming a stale pipe.
For example, resume the stale stalepipe1 pipe in the mydb.myschema database and schema:
select sys-tem$pipe_force_resume('mydb.myschema.stalepipe3','staleness_check_override'); As an event notification received while a pipe is paused reaches the end of the limited retention pe-riod, Snowflake schedules it to be dropped from the internal metadata. If the pipe is later resumed, Snowpipe processes these older notifications on a best effort basis. Snowflake cannot guarantee that they are processed.
For example, if a pipe is resumed 15 days after it was paused, Snowpipe generally skips any event notifications that were received on the first day the pipe was paused (i.e. that are now more than 14 days old).
If the pipe is resumed 16 days after it was paused, Snowpipe generally skips any event notifications that were received on the first and second days after the pipe was paused. And so on.


質問 # 26
A retail company stores transactions, store locations, and customer information tables in four reserved ra3.4xlarge Amazon Redshift cluster nodes. All three tables use even table distribution.
The company updates the store location table only once or twice every few years.
A data engineer notices that Redshift queues are slowing down because the whole store location table is constantly being broadcast to all four compute nodes for most queries. The data engineer wants to speed up the query performance by minimizing the broadcasting of the store location table.
Which solution will meet these requirements in the MOST cost-effective way?

  • A. Upgrade the Redshift reserved node to a larger instance size in the same instance family.
  • B. Change the distribution style of the store location table to KEY distribution based on the column that has the highest dimension.
  • C. Add a join column named store_id into the sort key for all the tables.
  • D. Change the distribution style of the store location table from EVEN distribution to ALL distribution.

正解:D

解説:
Changing the distribution style of the store location table to ALL distribution (A) is the most cost- effective solution. It directly addresses the issue of broadcasting by ensuring the entire table is available on each node, significantly improving join performance without incurring substantial additional costs.


質問 # 27
A company uses an Amazon Redshift cluster that runs on RA3 nodes. The company wants to scale read and write capacity to meet demand. A data engineer needs to identify a solution that will turn on concurrency scaling.
Which solution will meet this requirement?

  • A. Turn on concurrency scaling at the workload management (WLM) queue level in the Redshift cluster.
  • B. Turn on concurrency scaling in workload management (WLM) for Redshift Serverless workgroups.
  • C. Turn on concurrency scaling in the settings during the creation of any new Redshift cluster.
  • D. Turn on concurrency scaling for the daily usage quota for the Redshift cluster.

正解:A

解説:
https://docs.aws.amazon.com/redshift/latest/dg/concurrency-scaling-queues.html


質問 # 28
A marketing company uses Amazon S3 to store clickstream data. The company queries the data at the end of each day by using a SQL JOIN clause on S3 objects that are stored in separate buckets.
The company creates key performance indicators (KPIs) based on the objects. The company needs a serverless solution that will give users the ability to query data by partitioning the data.
The solution must maintain the atomicity, consistency, isolation, and durability (ACID) properties of the data.
Which solution will meet these requirements MOST cost-effectively?

  • A. Amazon S3 Select
  • B. Amazon EMR
  • C. Amazon Athena
  • D. Amazon Redshift Spectrum

正解:C


質問 # 29
Mark the Correct Statements:
Statement 1. Snowflake's zero-copy cloning feature provides a convenient way to quickly take a "snapshot" of any table, schema, or database.
Statement 2. Data Engineer can use zero-copy cloning feature for creating instant backups that do not incur any additional costs (until changes are made to the cloned object).

  • A. Statement 1
  • B. Statement 2
  • C. Statement 1 & 2 are correct.
  • D. Both are False.

正解:D

解説:
Explanation
Snowflake's zero-copy cloning feature provides a convenient way to quickly take a "snapshot" of any table, schema, or database and create a derived copy of that object which initially shares the underlying storage. This can be extremely useful for creating instant backups that do not incur any additional costs (until changes are made to the cloned object).
For example, when a clone is created of a table, the clone utilizes no data storage because it shares all the existing micro-partitions of the original table at the time it was cloned; however, rows can then be added, deleted, or updated in the clone independently from the original table. Each change to the clone results in new micro-partitions that are owned exclusively by the clone and are protect-ed through CDP.


質問 # 30
In a data engineering pipeline, a company is using multiple applications and teams to access a shared Amazon S3 bucket. To streamline access and simplify permissions management for these different entities, which S3 feature should the company utilize?

  • A. Enable multiple IAM roles, each corresponding to an application or team, granting access to the S3 bucket.
  • B. Use S3 Access Points to create unique endpoints with tailored permissions for each application or team.
  • C. Activate S3 Transfer Acceleration for the bucket to ensure fast and differentiated access for each application or team.
  • D. Implement S3 Lifecycle policies for each application or team to manage their specific data access and retention.

正解:B


質問 # 31
For enabling non-ACCOUNTADMIN Roles to Perform Data Sharing Tasks, which two glob-al/account privileges snowflake provide?

  • A. CREATE SHARE
  • B. REFERENCE USAGE
  • C. IMPORT SHARE
  • D. OPERATE

正解:A、C

解説:
Explanation
CREATE SHARE
In a provider account, this privilege enables creating and managing shares (for sharing data with consumer accounts).
IMPORT SHARE
In a consumer account, this privilege enables viewing the inbound shares shared with the account. Also enables creating databases from inbound shares; requires the global CREATE DATABASE privilege.
By default, these privileges are granted only to the ACCOUNTADMIN role, ensuring that only ac-count administrators can perform these tasks. However, the privileges can be granted to other roles, enabling the tasks to be delegated to other users in the account.


質問 # 32
A lab uses IoT sensors to monitor humidity, temperature, and pressure for a project. The sensors send 100 KB of data every 10 seconds. A downstream process will read the data from an Amazon S3 bucket every 30 seconds.
Which solution will deliver the data to the S3 bucket with the LEAST latency?

  • A. Use Amazon Managed Service for Apache Flink (previously known as Amazon Kinesis Data Analytics) and Amazon Kinesis Data Firehose to deliver the data to the S3 bucket. Use a 5 second buffer interval for Kinesis Data Firehose.
  • B. Use Amazon Kinesis Data Streams and Amazon Kinesis Data Firehose to deliver the data to the S3 bucket. Use the default buffer interval for Kinesis Data Firehose.
  • C. Use Amazon Kinesis Data Streams to deliver the data to the S3 bucket. Configure the stream to use 5 provisioned shards.
  • D. Use Amazon Kinesis Data Streams and call the Kinesis Client Library to deliver the data to the S3 bucket. Use a 5 second buffer interval from an application.

正解:A


質問 # 33
Ryan, a Data Engineer, accidently drop the Share named SF_SHARE which results in immediate access revoke for all the consumers (i.e., accounts who have created a database from that SF_SHARE). What action he can take to recover the dropped Share?

  • A. He can recreate a share with the same name as a previous share which does restore the databases created (by any consumers) from the share SF_SHARE.
  • B. By Executing UNDROP command he could possibly recover the dropped Share SF_SHARE & its associated Databases for immediate consumer access.
  • C. Consumer accounts that have created databases from the share will still be able to que-ry these databases as Share is separate securable object & it's still accessible using time travel feature.
  • D. A dropped share cannot be restored. The share must be created again using the CRE-ATE SHARE command and then configured using GRANT <privilege> ... TO SHARE and ALTER SHARE.

正解:D

解説:
Explanation
You can drop a share at any time using the DROP SHARE command.
Dropping a share instantly invalidates all databases created from the share by consumer accounts.
All queries and other operations performed on these databases will no longer work.
After dropping a share, you can recreate it with the same name;
however, this does not restore any of the databases created from the share by consumer accounts.
The recreated share is treated as a new share and all consumer accounts must create a new database from the new share.


質問 # 34
Elon, a Data Engineer, needs to Split Semi-structured Elements from the Source files and load them as an array into Separate Columns.
Source File:
1.+----------------------------------------------------------------------+
2.| $1 |
3.|----------------------------------------------------------------------|
4.| {"mac_address": {"host1": "197.128.1.1","host2": "197.168.0.1"}}, |
5.| {"mac_address": {"host1": "197.168.2.1","host2": "197.168.3.1"}} |
6.+----------------------------------------------------------------------+ Output: Splitting the Machine Address as below.
1.COL1 | COL2 |
2.|----------+----------|
3.| [ | [ |
4.| "197", | "197", |
5.| "128", | "168", |
6.| "1", | "0", |
7.| "1" | "1" |
8.| ] | ] |
9.| [ | [ |
10.| "197", | "197", |
11.| "168", | "168", |
12.| "2", | "3", |
13.| "1" | "1" |
14.| ] | ]
Which SnowFlake Function can Elon use to transform this semi structured data in the output for-mat?

  • A. SPLIT
  • B. NEST
  • C. GROUP_BY_CONNECT
  • D. CONVERT_TO_ARRAY

正解:A


質問 # 35
......

無料提供中のDEA-C01試験問題集で(2024年最新のPDF問題集)信頼度の高いDEA-C01テストエンジン:https://www.passtest.jp/Snowflake/DEA-C01-shiken.html

DEA-C01のPDFで最近更新された問題です集試験点数を伸ばそう:https://drive.google.com/open?id=1dp3xvcP1c3WU5P5MRC3-4WCb0LRNaEFR