Google Cloud Certified Professional-Data-Engineerリアル試験問題と無料最新回答2025年03月14日 [Q164-Q180]

Share

Google Cloud Certified Professional-Data-Engineerリアル試験問題と無料最新回答2025年03月14日

Professional-Data-Engineer究極な学習ガイド

質問 # 164
Which is not a valid reason for poor Cloud Bigtable performance?

  • A. The table's schema is not designed correctly.
  • B. The Cloud Bigtable cluster has too many nodes.
  • C. The workload isn't appropriate for Cloud Bigtable.
  • D. There are issues with the network connection.

正解:B

解説:
The Cloud Bigtable cluster doesn't have enough nodes. If your Cloud Bigtable cluster is overloaded, adding more nodes can improve performance. Use the monitoring tools to check whether the cluster is overloaded.


質問 # 165
Business owners at your company have given you a database of bank transactions. Each row contains the
user ID, transaction type, transaction location, and transaction amount. They ask you to investigate what
type of machine learning can be applied to the data. Which three machine learning applications can you
use? (Choose three.)

  • A. Unsupervised learning to determine which transactions are most likely to be fraudulent.
  • B. Reinforcement learning to predict the location of a transaction.
  • C. Clustering to divide the transactions into N categories based on feature similarity.
  • D. Unsupervised learning to predict the location of a transaction.
  • E. Supervised learning to determine which transactions are most likely to be fraudulent.
  • F. Supervised learning to predict the location of a transaction.

正解:A、B、C


質問 # 166
You work for a car manufacturer and have set up a data pipeline using Google Cloud Pub/Sub to capture anomalous sensor events. You are using a push subscription in Cloud Pub/Sub that calls a custom HTTPS endpoint that you have created to take action of these anomalous events as they occur. Your custom HTTPS endpoint keeps getting an inordinate amount of duplicate messages. What is the most likely cause of these duplicate messages?

  • A. The Cloud Pub/Sub topic has too many messages published to it.
  • B. The message body for the sensor event is too large.
  • C. Your custom endpoint is not acknowledging messages within the acknowledgement deadline.
  • D. Your custom endpoint has an out-of-date SSL certificate.

正解:C

解説:
Until or unless the message is not acknowledged within defined ack window period for every message, we will get duplicate (number of retries to send message can be defined).
https://cloud.google.com/pubsub/docs/troubleshooting#dupes


質問 # 167
Different teams in your organization store customer and performance data in BigOuery. Each team needs to keep full control of their collected data, be able to query data within their projects, and be able to exchange their data with other teams. You need to implement an organization-wide solution, while minimizing operational tasks and costs. What should you do?

  • A. Create a BigQuery scheduled query to replicate all customer data into team projects.
  • B. Ask each team to create authorized views of their data. Grant the biquery. jobUser role to each team.
  • C. Ask each team to publish their data in Analytics Hub. Direct the other teams to subscribe to them.
  • D. Enable each team to create materialized views of the data they need to access in their projects.

正解:C

解説:
To enable different teams to manage their own data while allowing data exchange across the organization, using Analytics Hub is the best approach. Here's why option C is the best choice:
Analytics Hub:
Analytics Hub allows teams to publish their data as data exchanges, making it easy for other teams to discover and subscribe to the data they need.
This approach maintains each team's control over their data while facilitating easy and secure data sharing across the organization.
Data Publishing and Subscribing:
Teams can publish datasets they control, allowing them to manage access and updates independently.
Other teams can subscribe to these published datasets, ensuring they have access to the latest data without duplicating efforts.
Minimized Operational Tasks and Costs:
This method reduces the need for complex replication or data synchronization processes, minimizing operational overhead.
By centralizing data sharing through Analytics Hub, it also reduces storage costs associated with duplicating large datasets.
Steps to Implement:
Set Up Analytics Hub:
Enable Analytics Hub in your Google Cloud project.
Provide training to teams on how to publish and subscribe to data exchanges.
Publish Data:
Each team publishes their datasets in Analytics Hub, configuring access controls and metadata as needed.
Subscribe to Data:
Teams that need access to data from other teams can subscribe to the relevant data exchanges, ensuring they always have up-to-date data.
Reference:
Analytics Hub Documentation
Publishing Data in Analytics Hub
Subscribing to Data in Analytics Hub


質問 # 168
Which software libraries are supported by Cloud Machine Learning Engine?

  • A. TensorFlow and Torch
  • B. TensorFlow
  • C. Theano and TensorFlow
  • D. Theano and Torch

正解:B

解説:
Cloud ML Engine mainly does two things:
Enables you to train machine learning models at scale by running TensorFlow training applications in the cloud.
Hosts those trained models for you in the cloud so that you can use them to get predictions about new data.


質問 # 169
Which is the preferred method to use to avoid hotspotting in time series data in Bigtable?

  • A. Field promotion
  • B. Hashing
  • C. Randomization
  • D. Salting

正解:A

解説:
By default, prefer field promotion. Field promotion avoids hotspotting in almost all cases, and it tends to make it easier to design a row key that facilitates queries.


質問 # 170
Google Cloud Bigtable indexes a single value in each row. This value is called the _______.

  • A. primary key
  • B. unique key
  • C. master key
  • D. row key

正解:D

解説:
Cloud Bigtable is a sparsely populated table that can scale to billions of rows and thousands of columns, allowing you to store terabytes or even petabytes of data. A single value in each row is indexed; this value is known as the row key.
Reference: https://cloud.google.com/bigtable/docs/overview


質問 # 171
You've migrated a Hadoop job from an on-prem cluster to dataproc and GCS. Your Spark job is a complicated analytical workload that consists of many shuffing operations and initial data are parquet files (on average 200-400 MB size each). You see some degradation in performance after the migration to Dataproc, so you'd like to optimize for it. You need to keep in mind that your organization is very cost-sensitive, so you'd like to continue using Dataproc on preemptibles (with 2 non-preemptible workers only) for this workload.
What should you do?

  • A. Increase the size of your parquet files to ensure them to be 1 GB minimum.
  • B. Switch from HDDs to SSDs, copy initial data from GCS to HDFS, run the Spark job and copy results back to GCS.
  • C. Switch from HDDs to SSDs, override the preemptible VMs configuration to increase the boot disk size.
  • D. Switch to TFRecords formats (appr. 200MB per file) instead of parquet files.

正解:A


質問 # 172
Each analytics team in your organization is running BigQuery jobs in their own projects. You want to enable each team to monitor slot usage within their projects. What should you do?

  • A. Create a Stackdriver Monitoring dashboard based on the BigQuery metric query/scanned_bytes
  • B. Create an aggregated log export at the organization level, capture the BigQuery job execution logs, create a custom metric based on the totalSlotMs, and create a Stackdriver Monitoring dashboard based on the custom metric
  • C. Create a log export for each project, capture the BigQuery job execution logs, create a custom metric based on the totalSlotMs, and create a Stackdriver Monitoring dashboard based on the custom metric
  • D. Create a Stackdriver Monitoring dashboard based on the BigQuery metric slots/allocated_for_project

正解:B


質問 # 173
You are using Cloud Bigtable to persist and serve stock market data for each of the major indices. To serve the trading application, you need to access only the most recent stock prices that are streaming in How should you design your row key and tables to ensure that you can access the data with the most simple query?

  • A. Create one unique table for all of the indices, and then use a reverse timestamp as the row key design.
  • B. Create one unique table for all of the indices, and then use the index and timestamp as the row key design
  • C. For each index, have a separate table and use a timestamp as the row key design
  • D. For each index, have a separate table and use a reverse timestamp as the row key design

正解:B


質問 # 174
You work for a manufacturing company that sources up to 750 different components, each from a different supplier. You've collected a labeled dataset that has on average 1000 examples for each unique component. Your team wants to implement an app to help warehouse workers recognize incoming components based on a photo of the component. You want to implement the first working version of this app (as Proof-Of-Concept) within a few working days. What should you do?

  • A. Use Cloud Vision AutoML, but reduce your dataset twice.
  • B. Use Cloud Vision API by providing custom labels as recognition hints.
  • C. Train your own image recognition model leveraging transfer learning techniques.
  • D. Use Cloud Vision AutoML with the existing dataset.

正解:D


質問 # 175
Your software uses a simple JSON format for all messages. These messages are published to Google Cloud Pub/Sub, then processed with Google Cloud Dataflow to create a real-time dashboard for the CFO.
During testing, you notice that some messages are missing in the dashboard. You check the logs, and all messages are being published to Cloud Pub/Sub successfully. What should you do next?

  • A. Use Google Stackdriver Monitoring on Cloud Pub/Sub to find the missing messages.
  • B. Switch Cloud Dataflow to pull messages from Cloud Pub/Sub instead of Cloud Pub/Sub pushing messages to Cloud Dataflow.
  • C. Run a fixed dataset through the Cloud Dataflow pipeline and analyze the output.
  • D. Check the dashboard application to see if it is not displaying correctly.

正解:A

解説:
Stackdriver can be used to check the error like number of unack messages, publisher pushing messages faster.


質問 # 176
You are operating a Cloud Dataflow streaming pipeline. The pipeline aggregates events from a Cloud Pub/ Sub subscription source, within a window, and sinks the resulting aggregation to a Cloud Storage bucket.
The source has consistent throughput. You want to monitor an alert on behavior of the pipeline with Cloud Stackdriver to ensure that it is processing data. Which Stackdriver alerts should you create?

  • A. An alert based on an increase of instance/storage/used_bytesfor the source and a rate of change decrease of subscription/num_undelivered_messages for the destination
  • B. An alert based on an increase of subscription/num_undelivered_messagesfor the source and a rate of change decrease of instance/storage/used_bytesfor the destination
  • C. An alert based on a decrease of instance/storage/used_bytesfor the source and a rate of change increase of subscription/num_undelivered_messages for the destination
  • D. An alert based on a decrease of subscription/num_undelivered_messagesfor the source and a rate of change increase of instance/storage/used_bytesfor the destination

正解:B


質問 # 177
Which of the following is not true about Dataflow pipelines?

  • A. Pipelines can share data between instances
  • B. Pipelines represent a directed graph of steps
  • C. Pipelines are a set of operations
  • D. Pipelines represent a data processing job

正解:A

解説:
Explanation
The data and transforms in a pipeline are unique to, and owned by, that pipeline. While your program can create multiple pipelines, pipelines cannot share data or transforms Reference: https://cloud.google.com/dataflow/model/pipelines


質問 # 178
Flowlogistic Case Study
Company Overview
Flowlogistic is a leading logistics and supply chain provider. They help businesses throughout the world manage their resources and transport them to their final destination. The company has grown rapidly, expanding their offerings to include rail, truck, aircraft, and oceanic shipping.
Company Background
The company started as a regional trucking company, and then expanded into other logistics market.
Because they have not updated their infrastructure, managing and tracking orders and shipments has become a bottleneck. To improve operations, Flowlogistic developed proprietary technology for tracking shipments in real time at the parcel level. However, they are unable to deploy it because their technology stack, based on Apache Kafka, cannot support the processing volume. In addition, Flowlogistic wants to further analyze their orders and shipments to determine how best to deploy their resources.
Solution Concept
Flowlogistic wants to implement two concepts using the cloud:
Use their proprietary technology in a real-time inventory-tracking system that indicates the location of

their loads
Perform analytics on all their orders and shipment logs, which contain both structured and unstructured

data, to determine how best to deploy resources, which markets to expand info. They also want to use predictive analytics to learn earlier when a shipment will be delayed.
Existing Technical Environment
Flowlogistic architecture resides in a single data center:
Databases

8 physical servers in 2 clusters
- SQL Server - user data, inventory, static data
3 physical servers
- Cassandra - metadata, tracking messages
10 Kafka servers - tracking message aggregation and batch insert
Application servers - customer front end, middleware for order/customs

60 virtual machines across 20 physical servers
- Tomcat - Java services
- Nginx - static content
- Batch servers
Storage appliances

- iSCSI for virtual machine (VM) hosts
- Fibre Channel storage area network (FC SAN) - SQL server storage
- Network-attached storage (NAS) image storage, logs, backups
10 Apache Hadoop /Spark servers

- Core Data Lake
- Data analysis workloads
20 miscellaneous servers

- Jenkins, monitoring, bastion hosts,
Business Requirements
Build a reliable and reproducible environment with scaled panty of production.

Aggregate data in a centralized Data Lake for analysis

Use historical data to perform predictive analytics on future shipments

Accurately track every shipment worldwide using proprietary technology

Improve business agility and speed of innovation through rapid provisioning of new resources

Analyze and optimize architecture for performance in the cloud

Migrate fully to the cloud if all other requirements are met

Technical Requirements
Handle both streaming and batch data

Migrate existing Hadoop workloads

Ensure architecture is scalable and elastic to meet the changing demands of the company.

Use managed services whenever possible

Encrypt data flight and at rest

Connect a VPN between the production data center and cloud environment

SEO Statement
We have grown so quickly that our inability to upgrade our infrastructure is really hampering further growth and efficiency. We are efficient at moving shipments around the world, but we are inefficient at moving data around.
We need to organize our information so we can more easily understand where our customers are and what they are shipping.
CTO Statement
IT has never been a priority for us, so as our data has grown, we have not invested enough in our technology. I have a good staff to manage IT, but they are so busy managing our infrastructure that I cannot get them to do the things that really matter, such as organizing our data, building the analytics, and figuring out how to implement the CFO' s tracking technology.
CFO Statement
Part of our competitive advantage is that we penalize ourselves for late shipments and deliveries. Knowing where out shipments are at all times has a direct correlation to our bottom line and profitability.
Additionally, I don't want to commit capital to building out a server environment.
Flowlogistic is rolling out their real-time inventory tracking system. The tracking devices will all send package-tracking messages, which will now go to a single Google Cloud Pub/Sub topic instead of the Apache Kafka cluster. A subscriber application will then process the messages for real-time reporting and store them in Google BigQuery for historical analysis. You want to ensure the package data can be analyzed over time.
Which approach should you take?

  • A. Use the NOW () function in BigQuery to record the event's time.
  • B. Use the automatically generated timestamp from Cloud Pub/Sub to order the data.
  • C. Attach the timestamp and Package ID on the outbound message from each publisher device as they are sent to Clod Pub/Sub.
  • D. Attach the timestamp on each message in the Cloud Pub/Sub subscriber application as they are received.

正解:C


質問 # 179
Why do you need to split a machine learning dataset into training data and test data?

  • A. To make sure your model is generalized for more than just the training data
  • B. So you can try two different sets of features
  • C. To allow you to create unit tests in your code
  • D. So you can use one dataset for a wide model and one for a deep model

正解:A

解説:
The flaw with evaluating a predictive model on training data is that it does not inform you on how well the model has generalized to new unseen data. A model that is selected for its accuracy on the training dataset rather than its accuracy on an unseen test dataset is very likely to have lower accuracy on an unseen test dataset. The reason is that the model is not as generalized. It has specialized to the structure in the training dataset. This is called overfitting.
Reference: https://machinelearningmastery.com/a-simple-intuition-for-overfitting/


質問 # 180
......

究極なガイド準備Professional-Data-Engineer認定試験Google Cloud Certified:https://www.passtest.jp/Google/Professional-Data-Engineer-shiken.html

リアルProfessional-Data-Engineer問題集でGoogle明確な解答を試そう:https://drive.google.com/open?id=1u5R4GxfkPXqb53oQvD6452YqPCteINhJ