[2025年01月] 合格させるGoogle Professional-Data-EngineerテストエンジンPDFで完全版無料問題集
Google Certified Professional Data Engineer Exam練習テスト2025年最新のProfessional-Data-Engineerストレスなしで合格!
Google Professional-Data-Enginer認定試験は、Google Cloudが提供する試験であり、Google Cloudプラットフォームでデータ処理システムの設計と構築に習熟する機会を個人に提供します。この認定は、データ処理の経験があり、Google Cloud Platform Technologiesと協力してきた個人を対象としています。認定試験では、データ処理システムを設計、構築、運用、安全、および監視する候補者の能力を測定します。
GoogleのProfessional-Data-Engineer認定試験は、プロフェッショナルなデータエンジニアとして自己を確立したい人々向けに、Googleが提供する高名な認定プログラムです。この認定は、データ処理システムの設計、構築、オペレーショナル化、セキュリティ、監視に必要なスキルや知識を検証します。データ処理システム、データウェアハウジング、データ分析技術での経験がある人々を対象としています。
質問 # 18
The Dataflow SDKs have been recently transitioned into which Apache service?
- A. Apache Hadoop
- B. Apache Kafka
- C. Apache Spark
- D. Apache Beam
正解:D
解説:
Explanation
Dataflow SDKs are being transitioned to Apache Beam, as per the latest Google directive Reference: https://cloud.google.com/dataflow/docs/
質問 # 19
Case Study: 2 - MJTelco
Company Overview
MJTelco is a startup that plans to build networks in rapidly growing, underserved markets around the world. The company has patents for innovative optical communications hardware. Based on these patents, they can create many reliable, high-speed backbone links with inexpensive hardware.
Company Background
Founded by experienced telecom executives, MJTelco uses technologies originally developed to overcome communications challenges in space. Fundamental to their operation, they need to create a distributed data infrastructure that drives real-time analysis and incorporates machine learning to continuously optimize their topologies. Because their hardware is inexpensive, they plan to overdeploy the network allowing them to account for the impact of dynamic regional politics on location availability and cost. Their management and operations teams are situated all around the globe creating many-to- many relationship between data consumers and provides in their system. After careful consideration, they decided public cloud is the perfect environment to support their needs.
Solution Concept
MJTelco is running a successful proof-of-concept (PoC) project in its labs. They have two primary needs:
Scale and harden their PoC to support significantly more data flows generated when they ramp to more than 50,000 installations.
Refine their machine-learning cycles to verify and improve the dynamic models they use to control topology definition.
MJTelco will also use three separate operating environments ?development/test, staging, and production ?
to meet the needs of running experiments, deploying new features, and serving production customers.
Business Requirements
Scale up their production environment with minimal cost, instantiating resources when and where needed in an unpredictable, distributed telecom user community. Ensure security of their proprietary data to protect their leading-edge machine learning and analysis.
Provide reliable and timely access to data for analysis from distributed research workers Maintain isolated environments that support rapid iteration of their machine-learning models without affecting their customers.
Technical Requirements
Ensure secure and efficient transport and storage of telemetry data Rapidly scale instances to support between 10,000 and 100,000 data providers with multiple flows each.
Allow analysis and presentation against data tables tracking up to 2 years of data storing approximately
100m records/day
Support rapid iteration of monitoring infrastructure focused on awareness of data pipeline problems both in telemetry flows and in production learning cycles.
CEO Statement
Our business model relies on our patents, analytics and dynamic machine learning. Our inexpensive hardware is organized to be highly reliable, which gives us cost advantages. We need to quickly stabilize our large distributed data pipelines to meet our reliability and capacity commitments.
CTO Statement
Our public cloud services must operate as advertised. We need resources that scale and keep our data secure. We also need environments in which our data scientists can carefully study and quickly adapt our models. Because we rely on automation to process our data, we also need our development and test environments to work as we iterate.
CFO Statement
The project is too large for us to maintain the hardware and software required for the data and analysis.
Also, we cannot afford to staff an operations team to monitor so many data feeds, so we will rely on automation and infrastructure. Google Cloud's machine learning will allow our quantitative researchers to work on our high-value problems instead of problems with our data pipelines.
MJTelco is building a custom interface to share data. They have these requirements:
They need to do aggregations over their petabyte-scale datasets. They need to scan specific time range rows with a very fast response time (milliseconds). Which combination of Google Cloud Platform products should you recommend?
- A. BigQuery and Cloud Storage
- B. BigQuery and Cloud Bigtable
- C. Cloud Datastore and Cloud Bigtable
- D. Cloud Bigtable and Cloud SQL
正解:B
質問 # 20
When you design a Google Cloud Bigtable schema it is recommended that you _________.
- A. Avoid schema designs that are based on NoSQL concepts
- B. Create schema designs that require atomicity across rows
- C. Create schema designs that are based on a relational database design
- D. Avoid schema designs that require atomicity across rows
正解:D
解説:
All operations are atomic at the row level. For example, if you update two rows in a table, it's possible that one row will be updated successfully and the other update will fail. Avoid schema designs that require atomicity across rows.
質問 # 21
Flowlogistic Case Study
Company Overview
Flowlogistic is a leading logistics and supply chain provider. They help businesses throughout the world manage their resources and transport them to their final destination. The company has grown rapidly, expanding their offerings to include rail, truck, aircraft, and oceanic shipping.
Company Background
The company started as a regional trucking company, and then expanded into other logistics market.
Because they have not updated their infrastructure, managing and tracking orders and shipments has become a bottleneck. To improve operations, Flowlogistic developed proprietary technology for tracking shipments in real time at the parcel level. However, they are unable to deploy it because their technology stack, based on Apache Kafka, cannot support the processing volume. In addition, Flowlogistic wants to further analyze their orders and shipments to determine how best to deploy their resources.
Solution Concept
Flowlogistic wants to implement two concepts using the cloud:
Use their proprietary technology in a real-time inventory-tracking system that indicates the location of
their loads
Perform analytics on all their orders and shipment logs, which contain both structured and unstructured
data, to determine how best to deploy resources, which markets to expand info. They also want to use predictive analytics to learn earlier when a shipment will be delayed.
Existing Technical Environment
Flowlogistic architecture resides in a single data center:
Databases
8 physical servers in 2 clusters
- SQL Server - user data, inventory, static data
3 physical servers
- Cassandra - metadata, tracking messages
10 Kafka servers - tracking message aggregation and batch insert
Application servers - customer front end, middleware for order/customs
60 virtual machines across 20 physical servers
- Tomcat - Java services
- Nginx - static content
- Batch servers
Storage appliances
- iSCSI for virtual machine (VM) hosts
- Fibre Channel storage area network (FC SAN) - SQL server storage
- Network-attached storage (NAS) image storage, logs, backups
Apache Hadoop /Spark servers
- Core Data Lake
- Data analysis workloads
20 miscellaneous servers
- Jenkins, monitoring, bastion hosts,
Business Requirements
Build a reliable and reproducible environment with scaled panty of production.
Aggregate data in a centralized Data Lake for analysis
Use historical data to perform predictive analytics on future shipments
Accurately track every shipment worldwide using proprietary technology
Improve business agility and speed of innovation through rapid provisioning of new resources
Analyze and optimize architecture for performance in the cloud
Migrate fully to the cloud if all other requirements are met
Technical Requirements
Handle both streaming and batch data
Migrate existing Hadoop workloads
Ensure architecture is scalable and elastic to meet the changing demands of the company.
Use managed services whenever possible
Encrypt data flight and at rest
Connect a VPN between the production data center and cloud environment
SEO Statement
We have grown so quickly that our inability to upgrade our infrastructure is really hampering further growth and efficiency. We are efficient at moving shipments around the world, but we are inefficient at moving data around.
We need to organize our information so we can more easily understand where our customers are and what they are shipping.
CTO Statement
IT has never been a priority for us, so as our data has grown, we have not invested enough in our technology. I have a good staff to manage IT, but they are so busy managing our infrastructure that I cannot get them to do the things that really matter, such as organizing our data, building the analytics, and figuring out how to implement the CFO' s tracking technology.
CFO Statement
Part of our competitive advantage is that we penalize ourselves for late shipments and deliveries. Knowing where out shipments are at all times has a direct correlation to our bottom line and profitability.
Additionally, I don't want to commit capital to building out a server environment.
Flowlogistic is rolling out their real-time inventory tracking system. The tracking devices will all send package-tracking messages, which will now go to a single Google Cloud Pub/Sub topic instead of the Apache Kafka cluster. A subscriber application will then process the messages for real-time reporting and store them in Google BigQuery for historical analysis. You want to ensure the package data can be analyzed over time.
Which approach should you take?
- A. Attach the timestamp and Package ID on the outbound message from each publisher device as they are sent to Clod Pub/Sub.
- B. Attach the timestamp on each message in the Cloud Pub/Sub subscriber application as they are received.
- C. Use the automatically generated timestamp from Cloud Pub/Sub to order the data.
- D. Use the NOW () function in BigQuery to record the event's time.
正解:A
質問 # 22
Your company receives both batch- and stream-based event data. You want to process the data using
Google Cloud Dataflow over a predictable time period. However, you realize that in some instances data
can arrive late or out of order. How should you design your Cloud Dataflow pipeline to handle data that is
late or out of order?
- A. Use watermarks and timestamps to capture the lagged data.
- B. Set a single global window to capture all the data.
- C. Ensure every datasource type (stream or batch) has a timestamp, and use the timestamps to define
the logic for lagged data. - D. Set sliding windows to capture all the lagged data.
正解:D
質問 # 23
Which Java SDK class can you use to run your Dataflow programs locally?
- A. DirectPipelineRunner
- B. MachineRunner
- C. LocalRunner
- D. LocalPipelineRunner
正解:A
解説:
DirectPipelineRunner allows you to execute operations in the pipeline directly, without any optimization. Useful for small local execution and tests Reference: https://cloud.google.com/dataflow/java- sdk/JavaDoc/com/google/cloud/dataflow/sdk/runners/DirectPipelineRunner
質問 # 24
Your company has hired a new data scientist who wants to perform complicated analyses across very large datasets stored in Google Cloud Storage and in a Cassandra cluster on Google Compute Engine.
The scientist primarily wants to create labelled data sets for machine learning projects, along with some visualization tasks. She reports that her laptop is not powerful enough to perform her tasks and it is slowing her down. You want to help her perform her tasks. What should you do?
- A. Grant the user access to Google Cloud Shell.
- B. Deploy Google Cloud Datalab to a virtual machine (VM) on Google Compute Engine.
- C. Host a visualization tool on a VM on Google Compute Engine.
- D. Run a local version of Jupiter on the laptop.
正解:B
解説:
Datalab provides Jupyter for this kind of work.
質問 # 25
Your company has a hybrid cloud initiative. You have a complex data pipeline that moves data between cloud provider services and leverages services from each of the cloud providers. Which cloud-native service should you use to orchestrate the entire pipeline?
- A. Cloud Dataflow
- B. Cloud Composer
- C. Cloud Dataprep
- D. Cloud Dataproc
正解:B
解説:
Cloud Composer uses airflow which is open source and can help to orchestrate jobs.
質問 # 26
You are building a data pipeline on Google Cloud. You need to prepare data using a casual method for a machine-learning process. You want to support a logistic regression model. You also need to monitor and adjust for null values, which must remain real-valued and cannot be removed. What should you do?
- A. Use Cloud Dataflow to find null values in sample source data. Convert all nulls to `none' using a Cloud Dataprep job.
- B. Use Cloud Dataprep to find null values in sample source data. Convert all nulls to 0 using a Cloud Dataprep job.
- C. Use Cloud Dataflow to find null values in sample source data. Convert all nulls to using a custom script.
- D. Use Cloud Dataprep to find null values in sample source data. Convert all nulls to `none' using a Cloud Dataproc job.
正解:B
質問 # 27
Your company is in a highly regulated industry. One of your requirements is to ensure individual users have access only to the minimum amount of information required to do their jobs. You want to enforce this requirement with Google BigQuery.
Which three approaches can you take? (Choose three.)
- A. Restrict BigQuery API access to approved users.
- B. Use Google Stackdriver Audit Logging to determine policy violations.
- C. Restrict access to tables by role.
- D. Disable writes to certain tables.
- E. Ensure that the data is encrypted at all times.
- F. Segregate data across multiple tables or databases.
正解:A、B、C
解説:
bigquery.tables.create Create new tables.
bigquery.tables.delete Delete tables.
bigquery.tables.export Export table data out of BigQuery.
bigquery.tables.get Get table metadata.
To get table data, you need bigquery.tables.getData.
bigquery.tables.getData Get table data. This permission is required for querying table data.
To get table metadata, you need bigquery.tables.get.
bigquery.tables.list List tables and metadata on tables.
bigquery.tables.setCategory Set policy tags in table schema.
bigquery.tables.update
Update table metadata.
To update table data, you need bigquery.tables.updateData.
bigquery.tables.updateData
Update table data.
To update table metadata, you need bigquery.tables.update.
質問 # 28
You are running a pipeline in Cloud Dataflow that receives messages from a Cloud Pub/Sub topic and writes the results to a BigQuery dataset in the EU. Currently, your pipeline is located in europe-west4 and has a maximum of 3 workers, instance type n1-standard-1. You notice that during peak periods, your pipeline is struggling to process records in a timely fashion, when all 3 workers are at maximum CPU utilization. Which two actions can you take to increase performance of your pipeline? (Choose two.)
- A. Increase the number of max workers
- B. Create a temporary table in Cloud Spanner that will act as a buffer for new data. Create a new step in your pipeline to write to this table first, and then create a new pipeline to write from Cloud Spanner to BigQuery
- C. Use a larger instance type for your Cloud Dataflow workers
- D. Change the zone of your Cloud Dataflow pipeline to run in us-central1
- E. Create a temporary table in Cloud Bigtable that will act as a buffer for new data. Create a new step in your pipeline to write to this table first, and then create a new pipeline to write from Cloud Bigtable to BigQuery
正解:B、C
質問 # 29
MJTelco needs you to create a schema in Google Bigtable that will allow for the historical analysis of the last
2 years of records. Each record that comes in is sent every 15 minutes, and contains a unique identifier of the device and a data record. The most common query is for all the data for a given device for a given day. Which schema should you use?
- A. Rowkey: date#data_pointColumn data: device_id
- B. Rowkey: data_pointColumn data: device_id, date
- C. Rowkey: date#device_idColumn data: data_point
- D. Rowkey: dateColumn data: device_id, data_point
- E. Rowkey: device_idColumn data: date, data_point
正解:B
質問 # 30
You need to choose a database to store time series CPU and memory usage for millions of computers. You need to store this data in one-second interval samples. Analysts will be performing real-time, ad hoc analytics against the database. You want to avoid being charged for every query executed and ensure that the schema design will allow for future growth of the dataset. Which database and data model should you choose?
- A. Create a table in BigQuery, and append the new samples for CPU and memory to the table
- B. Create a narrow table in Cloud Bigtable with a row key that combines the Computer Engine computer identifier with the sample time at each second
- C. Create a wide table in Cloud Bigtable with a row key that combines the computer identifier with the sample time at each minute, and combine the values for each second as column data.
- D. Create a wide table in BigQuery, create a column for the sample value at each second, and update the row with the interval for each second
正解:B
質問 # 31
You need to choose a database for a new project that has the following requirements:
* Fully managed
* Able to automatically scale up
* Transactionally consistent
* Able to scale up to 6 TB
* Able to be queried using SQL
Which database do you choose?
- A. Cloud Spanner
- B. Cloud Datastore
- C. Cloud Bigtable
- D. Cloud SQL
正解:A
質問 # 32
You are developing an application on Google Cloud that will automatically generate subject labels for users' blog posts. You are under competitive pressure to add this feature quickly, and you have no additional developer resources. No one on your team has experience with machine learning. What should you do?
- A. Build and train a text classification model using TensorFlow. Deploy the model using a Kubernetes Engine cluster. Call the model from your application and process the results as labels.
- B. Call the Cloud Natural Language API from your application. Process the generated Sentiment Analysis as labels.
- C. Call the Cloud Natural Language API from your application. Process the generated Entity Analysis as
labels. - D. Build and train a text classification model using TensorFlow. Deploy the model using Cloud Machine
Learning Engine. Call the model from your application and process the results as labels.
正解:B
質問 # 33
You work for a large bank that operates in locations throughout North America. You are setting up a data storage system that will handle bank account transactions. You require ACID compliance and the ability to access data with SQL. Which solution is appropriate?
- A. Store transaction data in Cloud Spanner. Enable stale reads to reduce latency.
- B. Store transaction data in BigQuery. Disabled the query cache to ensure consistency.
- C. Store transaction data in Cloud SQL. Use a federated query BigQuery for analysis.
- D. Store transaction in Cloud Spanner. Use locking read-write transactions.
正解:B
質問 # 34
You are choosing a NoSQL database to handle telemetry data submitted from millions of Internet-of- Things (IoT) devices. The volume of data is growing at 100 TB per year, and each data entry has about
100 attributes. The data processing pipeline does not require atomicity, consistency, isolation, and durability (ACID). However, high availability and low latency are required.
You need to analyze the data by querying against individual fields. Which three databases meet your requirements? (Choose three.)
- A. HDFS with Hive
- B. Redis
- C. MySQL
- D. HBase
- E. Cassandra
- F. MongoDB
正解:A、D、F
質問 # 35
You are building a model to make clothing recommendations. You know a user's fashion preference is likely to change over time, so you build a data pipeline to stream new data back to the model as it becomes available.
How should you use this data to train the model?
- A. Continuously retrain the model on just the new data.
- B. Train on the new data while using the existing data as your test set.
- C. Continuously retrain the model on a combination of existing data and the new data.
- D. Train on the existing data while using the new data as your test set.
正解:D
解説:
https://cloud.google.com/automl-tables/docs/prepare
質問 # 36
MJTelco's Google Cloud Dataflow pipeline is now ready to start receiving data from the 50,000 installations. You want to allow Cloud Dataflow to scale its compute power up as required. Which Cloud Dataflow pipeline configuration setting should you update?
- A. The zone
- B. The maximum number of workers
- C. The disk size per worker
- D. The number of workers
正解:A
質問 # 37
Your company is streaming real-time sensor data from their factory floor into Bigtable and they have
noticed extremely poor performance. How should the row key be redesigned to improve Bigtable
performance on queries that populate real-time dashboards?
- A. Use a row key of the form >#<sensorid>#<timestamp>.
- B. Use a row key of the form <timestamp>#<sensorid>.
- C. Use a row key of the form <timestamp>.
- D. Use a row key of the form <sensorid>.
正解:C
質問 # 38
You are a retailer that wants to integrate your online sales capabilities with different in-home assistants, such as Google Home. You need to interpret customer voice commands and issue an order to the backend systems.
Which solutions should you choose?
- A. Cloud AutoML Natural Language
- B. Cloud Natural Language API
- C. Cloud Speech-to-Text API
- D. Dialogflow Enterprise Edition
正解:A
質問 # 39
......
Google Professional-Data-Engineer認定試験は、データエンジニアリングの分野の候補者の知識とスキルをテストするように設計されています。この試験は、データ処理システムの設計、構築、維持を担当する専門家を対象としています。この試験は、Google Cloud Platform Technologiesを使用してデータ処理システムを設計および実装し、データ構造とデータベースを構築および維持し、データ処理ワークフローを分析および最適化するために、Google Cloud Platform Technologiesを使用する候補者の能力を検証するように設計されています。
時間限定!今すぐ無料アクセスProfessional-Data-Engineer練習試験用問題:https://drive.google.com/open?id=1-fxsZpqZndadOGIFbRdNF_u31TuQSeml
オンライン試験練習テストと詳細な解説付き!:https://www.passtest.jp/Google/Professional-Data-Engineer-shiken.html