[2023年12月09日]Professional-Data-Engineer試験問題集でリアル試験と100%同じ問題と解答 [Q139-Q164]

Share

[2023年12月09日]Professional-Data-Engineer試験問題集でリアル試験と100%同じ問題と解答

Professional-Data-Engineerテストエンジン問題集トレーニングには270問あります


Google Professional-Data-Engineer試験は、Google Cloud Platformが提供するデータ専門家向けの認定試験であり、Google Cloud Platform上でデータ処理システムを設計、構築、および管理する能力を証明したい人々向けに設計されています。この試験は、業界で高く評価され、特にビッグデータを扱いたい人々にとって重要な認定試験です。この試験は、様々なデータエンジニアリングツールや技術に関する候補者の知識をテストし、試験に合格することで、候補者がGoogle Cloud Platform上でデータソリューションを設計・実装するためのスキルと知識を持っていることを証明します。


Google Professional-Data-Enginer(Google Certified Professional Data Engineer)試験は、データ処理システムの設計、構築、管理におけるデータエンジニアのスキルを検証する認定プログラムです。 Google Cloud Technologiesを使用して、スケーラブルで信頼性が高く、効率的なデータパイプラインを開発および展開するためのデータエンジニアの習熟度を評価するように設計されています。この試験では、データ処理アーキテクチャ、データモデリング、データ分析、機械学習など、さまざまなトピックをカバーしています。

 

質問 # 139
You are updating the code for a subscriber to a Put/Sub feed. You are concerned that upon deployment the subscriber may erroneously acknowledge messages, leading to message loss. You subscriber is not set up to retain acknowledged messages. What should you do to ensure that you can recover from errors after deployment?

  • A. Use Cloud Build for your deployment if an error occurs after deployment, use a Seek operation to locate a tmestamp logged by Cloud Build at the start of the deployment
  • B. Enable dead-lettering on the Pub/Sub topic to capture messages that aren't successful acknowledged if an error occurs after deployment, re-deliver any messages captured by the dead-letter queue
  • C. Create a Pub/Sub snapshot before deploying new subscriber code. Use a Seek operation to re-deliver messages that became available after the snapshot was created
  • D. Set up the Pub/Sub emulator on your local machine Validate the behavior of your new subscriber togs before deploying it to production

正解:C


質問 # 140
Cloud Dataproc charges you only for what you really use with _____ billing.

  • A. month-by-month
  • B. minute-by-minute
  • C. week-by-week
  • D. hour-by-hour

正解:B

解説:
Explanation
One of the advantages of Cloud Dataproc is its low cost. Dataproc charges for what you really use with minute-by-minute billing and a low, ten-minute-minimum billing period.
Reference: https://cloud.google.com/dataproc/docs/concepts/overview


質問 # 141
Your company is running their first dynamic campaign, serving different offers by analyzing real-time data
during the holiday season. The data scientists are collecting terabytes of data that rapidly grows every
hour during their 30-day campaign. They are using Google Cloud Dataflow to preprocess the data and
collect the feature (signals) data that is needed for the machine learning model in Google Cloud Bigtable.
The team is observing suboptimal performance with reads and writes of their initial load of 10 TB of data.
They want to improve this performance while minimizing cost. What should they do?

  • A. The performance issue should be resolved over time as the site of the BigDate cluster is increased.
  • B. Redefine the schema by evenly distributing reads and writes across the row space of the table.
  • C. Redesign the schema to use a single row key to identify values that need to be updated frequently in
    the cluster.
  • D. Redesign the schema to use row keys based on numeric IDs that increase sequentially per user
    viewing the offers.

正解:B


質問 # 142
Your company is implementing a data warehouse using BigQuery, and you have been tasked with designing the data model You move your on-premises sales data warehouse with a star data schema to BigQuery but notice performance issues when querying the data of the past 30 days Based on Google's recommended practices, what should you do to speed up the query without increasing storage costs?

  • A. Shard the data by customer ID
  • B. Materialize the dimensional data in views
  • C. Partition the data by transaction date
  • D. Denormalize the data

正解:B


質問 # 143
You need to compose visualization for operations teams with the following requirements:
Telemetry must include data from all 50,000 installations for the most recent 6 weeks (sampling once every minute)
The report must not be more than 3 hours delayed from live data.
The actionable report should only show suboptimal links.
Most suboptimal links should be sorted to the top.
Suboptimal links can be grouped and filtered by regional geography.
User response time to load the report must be <5 seconds.
You create a data source to store the last 6 weeks of data, and create visualizations that allow viewers to see multiple date ranges, distinct geographic regions, and unique installation types. You always show the latest data without any changes to your visualizations. You want to avoid creating and updating new visualizations each month. What should you do?

  • A. Look through the current data and compose a small set of generalized charts and tables bound to criteria filters that allow value selection.
  • B. Look through the current data and compose a series of charts and tables, one for each possible
    combination of criteria.
  • C. Load the data into relational database tables, write a Google App Engine application that queries all rows, summarizes the data across each criteria, and then renders results using the Google Charts and visualization API.
  • D. Export the data to a spreadsheet, compose a series of charts and tables, one for each possible
    combination of criteria, and spread them across multiple tabs.

正解:A


質問 # 144
You want to use Google Stackdriver Logging to monitor Google BigQuery usage. You need an instant notification to be sent to your monitoring tool when new data is appended to a certain table using an insert job, but you do not want to receive notifications for other tables. What should you do?

  • A. In the Stackdriver logging admin interface, and enable a log sink export to BigQuery.
  • B. In the Stackdriver logging admin interface, enable a log sink export to Google Cloud Pub/Sub, and subscribe to the topic from your monitoring tool.
  • C. Make a call to the Stackdriver API to list all logs, and apply an advanced filter.
  • D. Using the Stackdriver API, create a project sink with advanced log filter to export to Pub/Sub, and subscribe to the topic from your monitoring tool.

正解:D

解説:
A and B are wrong since don't notify anything to the monitoring tool.
C has no filter on what will be notified. We want only some tables.


質問 # 145
When you store data in Cloud Bigtable, what is the recommended minimum amount of stored data?

  • A. 500 GB
  • B. 1 GB
  • C. 1 TB
  • D. 500 TB

正解:C

解説:
Cloud Bigtable is not a relational database. It does not support SQL queries, joins, or multi- row transactions. It is not a good solution for less than 1 TB of data.
Reference:
https://cloud.google.com/bigtable/docs/overview#title_short_and_other_storage_options


質問 # 146
You are implementing several batch jobs that must be executed on a schedule. These jobs have many interdependent steps that must be executed in a specific order. Portions of the jobs involve executing shell scripts, running Hadoop jobs, and running queries in BigQuery. The jobs are expected to run for many minutes up to several hours. If the steps fail, they must be retried a fixed number of times. Which service should you use to manage the execution of these jobs?

  • A. Cloud Composer
  • B. Cloud Functions
  • C. Cloud Scheduler
  • D. Cloud Dataflow

正解:C


質問 # 147
You work for an economic consulting firm that helps companies identify economic trends as they happen. As part of your analysis, you use Google BigQuery to correlate customer data with the average prices of the 100 most common goods sold, including bread, gasoline, milk, and others. The average prices of these goods are updated every 30 minutes. You want to make sure this data stays up to date so you can combine it with other data in BigQuery as cheaply as possible. What should you do?

  • A. Store and update the data in a regional Google Cloud Storage bucket and create a federated data source in BigQuery
  • B. Store the data in a file in a regional Google Cloud Storage bucket. Use Cloud Dataflow to query BigQuery and combine the data programmatically with the data stored in Google Cloud Storage.
  • C. Load the data every 30 minutes into a new partitioned table in BigQuery.
  • D. Store the data in Google Cloud Datastore. Use Google Cloud Dataflow to query BigQuery and combine the data programmatically with the data stored in Cloud Datastore

正解:D


質問 # 148
You have enabled the free integration between Firebase Analytics and Google BigQuery. Firebase now
automatically creates a new table daily in BigQuery in the format app_events_YYYYMMDD.You want to
query all of the tables for the past 30 days in legacy SQL. What should you do?

  • A. Use the TABLE_DATE_RANGEfunction
  • B. Use WHEREdate BETWEEN YYYY-MM-DD AND YYYY-MM-DD
  • C. Use SELECT IF.(date >= YYYY-MM-DD AND date <= YYYY-MM-DD
  • D. Use the WHERE_PARTITIONTIMEpseudo column

正解:A

解説:
Explanation/Reference:
Reference: https://cloud.google.com/blog/products/gcp/using-bigquery-and-firebase-analytics-to-
understand-your-mobile-app?hl=am


質問 # 149
Your company produces 20,000 files every hour. Each data file is formatted as a comma separated values (CSV) file that is less than 4 KB. All files must be ingested on Google Cloud Platform before they can be processed. Your company site has a 200 ms latency to Google Cloud, and your Internet connection bandwidth is limited as 50 Mbps. You currently deploy a secure FTP (SFTP) server on a virtual machine in Google Compute Engine as the data ingestion point. A local SFTP client runs on a dedicated machine to transmit the CSV files as is. The goal is to make reports with data from the previous day available to the executives by
10:00 a.m. each day. This design is barely able to keep up with the current volume, even though the bandwidth utilization is rather low.
You are told that due to seasonality, your company expects the number of files to double for the next three months. Which two actions should you take? (choose two.)

  • A. Redesign the data ingestion process to use gsutil tool to send the CSV files to a storage bucket in parallel.
  • B. Create an S3-compatible storage endpoint in your network, and use Google Cloud Storage Transfer Service to transfer on-premises data to the designated storage bucket.
  • C. Introduce data compression for each file to increase the rate file of file transfer.
  • D. Contact your internet service provider (ISP) to increase your maximum bandwidth to at least 100 Mbps.
  • E. Assemble 1,000 files into a tape archive (TAR) file. Transmit the TAR files instead, and disassemble the CSV files in the cloud upon receiving them.

正解:A、B


質問 # 150
You want to archive data in Cloud Storage. Because some data is very sensitive, you want to use the
"Trust No One" (TNO) approach to encrypt your data to prevent the cloud provider staff from decrypting your data. What should you do?

  • A. Use gcloud kms keys create to create a symmetric key. Then use gcloud kms encrypt to encrypt each archival file with the key and unique additional authenticated data (AAD). Use gsutil to upload each encrypted file to the Cloud Storage bucket, and keep the AAD outside of Google cp Cloud.
  • B. Use gcloud kms keys create to create a symmetric key. Then use gcloud kms encrypt to encrypt each archival file with the key. Use gsutil cp to upload each encrypted file to the Cloud Storage bucket.
    Manually destroy the key previously used for encryption, and rotate the key once.
  • C. Specify customer-supplied encryption key (CSEK) in the .boto configuration file. Use gsutil cp to upload each archival file to the Cloud Storage bucket. Save the CSEK in Cloud Memorystore as permanent storage of the secret.
  • D. Specify customer-supplied encryption key (CSEK) in the .boto configuration file. Use gsutil cp to upload each archival file to the Cloud Storage bucket. Save the CSEK in a different project that only the security team can access.

正解:D


質問 # 151
You are designing a basket abandonment system for an ecommerce company. The system will send a message to a user based on these rules:
No interaction by the user on the site for 1 hour

Has added more than $30 worth of products to the basket Has not completed a

transaction
You use Google Cloud Dataflow to process the data and decide if a message should be sent. How should you design the pipeline?

  • A. Use a sliding time window with a duration of 60 minutes.
  • B. Use a global window with a time based trigger with a delay of 60 minutes.
  • C. Use a session window with a gap time duration of 60 minutes.
  • D. Use a fixed-time window with a duration of 60 minutes.

正解:B


質問 # 152
Case Study 2 - MJTelco
Company Overview
MJTelco is a startup that plans to build networks in rapidly growing, underserved markets around the world.
The company has patents for innovative optical communications hardware. Based on these patents, they can create many reliable, high-speed backbone links with inexpensive hardware.
Company Background
Founded by experienced telecom executives, MJTelco uses technologies originally developed to overcome communications challenges in space. Fundamental to their operation, they need to create a distributed data infrastructure that drives real-time analysis and incorporates machine learning to continuously optimize their topologies. Because their hardware is inexpensive, they plan to overdeploy the network allowing them to account for the impact of dynamic regional politics on location availability and cost.
Their management and operations teams are situated all around the globe creating many-to-many relationship between data consumers and provides in their system. After careful consideration, they decided public cloud is the perfect environment to support their needs.
Solution Concept
MJTelco is running a successful proof-of-concept (PoC) project in its labs. They have two primary needs:
* Scale and harden their PoC to support significantly more data flows generated when they ramp to more than 50,000 installations.
* Refine their machine-learning cycles to verify and improve the dynamic models they use to control topology definition.
MJTelco will also use three separate operating environments - development/test, staging, and production - to meet the needs of running experiments, deploying new features, and serving production customers.
Business Requirements
* Scale up their production environment with minimal cost, instantiating resources when and where needed in an unpredictable, distributed telecom user community.
* Ensure security of their proprietary data to protect their leading-edge machine learning and analysis.
* Provide reliable and timely access to data for analysis from distributed research workers
* Maintain isolated environments that support rapid iteration of their machine-learning models without affecting their customers.
Technical Requirements
* Ensure secure and efficient transport and storage of telemetry data
* Rapidly scale instances to support between 10,000 and 100,000 data providers with multiple flows each.
* Allow analysis and presentation against data tables tracking up to 2 years of data storing approximately
100m records/day
* Support rapid iteration of monitoring infrastructure focused on awareness of data pipeline problems both in telemetry flows and in production learning cycles.
CEO Statement
Our business model relies on our patents, analytics and dynamic machine learning. Our inexpensive hardware is organized to be highly reliable, which gives us cost advantages. We need to quickly stabilize our large distributed data pipelines to meet our reliability and capacity commitments.
CTO Statement
Our public cloud services must operate as advertised. We need resources that scale and keep our data secure. We also need environments in which our data scientists can carefully study and quickly adapt our models. Because we rely on automation to process our data, we also need our development and test environments to work as we iterate.
CFO Statement
The project is too large for us to maintain the hardware and software required for the data and analysis.
Also, we cannot afford to staff an operations team to monitor so many data feeds, so we will rely on automation and infrastructure. Google Cloud's machine learning will allow our quantitative researchers to work on our high-value problems instead of problems with our data pipelines.
Given the record streams MJTelco is interested in ingesting per day, they are concerned about the cost of Google BigQuery increasing. MJTelco asks you to provide a design solution. They require a single large data table called tracking_table. Additionally, they want to minimize the cost of daily queries while performing fine-grained analysis of each day's events. They also want to use streaming ingestion. What should you do?

  • A. Create a table called tracking_table and include a DATE column.
  • B. Create a partitioned table called tracking_table and include a TIMESTAMP column.
  • C. Create a table called tracking_table with a TIMESTAMP column to represent the day.
  • D. Create sharded tables for each day following the pattern tracking_table_YYYYMMDD.

正解:B


質問 # 153
You are a retailer that wants to integrate your online sales capabilities with different in-home assistants, such as Google Home. You need to interpret customer voice commands and issue an order to the backend systems. Which solutions should you choose?

  • A. Cloud Natural Language API
  • B. Cloud Speech-to-Text API
  • C. Dialogflow Enterprise Edition
  • D. Cloud AutoML Natural Language

正解:D


質問 # 154
You are choosing a NoSQL database to handle telemetry data submitted from millions of Internet-of-Things (IoT) devices. The volume of data is growing at 100 TB per year, and each data entry has about 100 attributes.
The data processing pipeline does not require atomicity, consistency, isolation, and durability (ACID).
However, high availability and low latency are required.
You need to analyze the data by querying against individual fields. Which three databases meet your requirements? (Choose three.)

  • A. Redis
  • B. MySQL
  • C. HBase
  • D. HDFS with Hive
  • E. MongoDB
  • F. Cassandra

正解:C、D、E


質問 # 155
Your neural network model is taking days to train. You want to increase the training speed. What can you
do?

  • A. Increase the number of layers in your neural network.
  • B. Subsample your test dataset.
  • C. Increase the number of input features to your model.
  • D. Subsample your training dataset.

正解:A

解説:
Explanation/Reference:
Reference: https://towardsdatascience.com/how-to-increase-the-accuracy-of-a-neural-network-
9f5d1c6f407d


質問 # 156
You are designing storage for very large text files for a data pipeline on Google Cloud. You want to support ANSI SQL queries. You also want to support compression and parallel load from the input locations using Google recommended practices. What should you do?

  • A. Compress text files to gzip using the Grid Computing Tools. Use BigQuery for storage and query.
  • B. Compress text files to gzip using the Grid Computing Tools. Use Cloud Storage, and then import into Cloud Bigtable for query.
  • C. Transform text files to compressed Avro using Cloud Dataflow. Use Cloud Storage and BigQuery permanent linked tables for query.
  • D. Transform text files to compressed Avro using Cloud Dataflow. Use BigQuery for storage and query.

正解:B

解説:
Explanation/Reference:


質問 # 157
Does Dataflow process batch data pipelines or streaming data pipelines?

  • A. None of the above
  • B. Both Batch and Streaming Data Pipelines
  • C. Only Batch Data Pipelines
  • D. Only Streaming Data Pipelines

正解:B

解説:
Dataflow is a unified processing model, and can execute both streaming and batch data pipelines


質問 # 158
Which of the following statements is NOT true regarding Bigtable access roles?

  • A. You can configure access control only at the project level.
  • B. Using IAM roles, you cannot give a user access to only one table in a project, rather than all tables in a project.
  • C. To give a user access to only one table in a project, grant the user the Bigtable Editor role for that table.
  • D. To give a user access to only one table in a project, you must configure access through your application.

正解:C

解説:
For Cloud Bigtable, you can configure access control at the project level. For example, you can grant the ability to:
Read from, but not write to, any table within the project. Read from and write to any table within the project, but not manage instances. Read from and write to any table within the project, and manage instances.
Reference: https://cloud.google.com/bigtable/docs/access-control


質問 # 159
You are designing a data processing pipeline. The pipeline must be able to scale automatically as load increases. Messages must be processed at least once, and must be ordered within windows of 1 hour. How should you design the solution?

  • A. Use Apache Kafka for message ingestion and use Cloud Dataflow for streaming analysis.
  • B. Use Cloud Pub/Sub for message ingestion and Cloud Dataflow for streaming analysis.
  • C. Use Apache Kafka for message ingestion and use Cloud Dataproc for streaming analysis.
  • D. Use Cloud Pub/Sub for message ingestion and Cloud Dataproc for streaming analysis.

正解:D


質問 # 160
As your organization expands its usage of GCP, many teams have started to create their own projects. Projects are further multiplied to accommodate different stages of deployments and target audiences. Each project requires unique access control configurations. The central IT team needs to have access to all projects. Furthermore, data from Cloud Storage buckets and BigQuery datasets must be shared for use in other projects in an ad hoc way. You want to simplify access control management by minimizing the number of policies. Which two steps should you take? Choose 2 answers.

  • A. Create distinct groups for various teams, and specify groups in Cloud IAM policies.
  • B. Use Cloud Deployment Manager to automate access provision.
  • C. Only use service accounts when sharing data for Cloud Storage buckets and BigQuery datasets.
  • D. Introduce resource hierarchy to leverage access control policy inheritance.
  • E. For each Cloud Storage bucket or BigQuery dataset, decide which projects need access. Find all the active members who have access to these projects, and create a Cloud IAM policy to grant access to all these users.

正解:A、B


質問 # 161
A shipping company has live package-tracking data that is sent to an Apache Kafka stream in real time. This is then loaded into BigQuery. Analysts in your company want to query the tracking data in BigQuery to analyze geospatial trends in the lifecycle of a package. The table was originally created with ingest-date partitioning.
Over time, the query processing time has increased. You need to implement a change that would improve query performance in BigQuery. What should you do?

  • A. Tier older data onto Google Cloud Storage files and create a BigQuery table using GCS as an external data source.
  • B. Implement clustering in BigQuery on the ingest date column.
  • C. Re-create the table using data partitioning on the package delivery date.
  • D. Implement clustering in BigQuery on the package-tracking ID column.

正解:D


質問 # 162
Which role must be assigned to a service account used by the virtual machines in a Dataproc cluster so they can execute jobs?

  • A. Dataproc Runner
  • B. Dataproc Worker
  • C. Dataproc Editor
  • D. Dataproc Viewer

正解:B

解説:
Service accounts used with Cloud Dataproc must have Dataproc/Dataproc Worker role (or have all the permissions granted by Dataproc Worker role).
Reference: https://cloud.google.com/dataproc/docs/concepts/service-accounts#important_notes


質問 # 163
Your company needs to upload their historic data to Cloud Storage. The security rules don't allow access from external IPs to their on-premises resources. After an initial upload, they will add new data from existing on-premises applications every day. What should they do?

  • A. Install an FTP server on a Compute Engine VM to receive the files and move them to Cloud Storage.
  • B. Use Cloud Dataflow and write the data to Cloud Storage.
  • C. Write a job template in Cloud Dataproc to perform the data transfer.
  • D. Execute gsutil rsyncfrom the on-premises servers.

正解:B


質問 # 164
......


Google Professional-Data-Engineer試験は、選択式および複数選択式の問題、およびシナリオベースの問題から構成され、候補者が問題解決スキルを示す必要がある。試験時間は3時間で、合格には少なくとも70%のスコアが必要です。試験料は200ドルで、英語、日本語、スペイン語で利用可能です。

 

Professional-Data-Engineer練習テストPDF試験材料:https://www.passtest.jp/Google/Professional-Data-Engineer-shiken.html

Professional-Data-Engineer問題で一発合格させる問題集にはGoogle Cloud Certified認定問題を使おう:https://drive.google.com/open?id=1-fxsZpqZndadOGIFbRdNF_u31TuQSeml