跳到正文
Semiconductor Engineering 芯片与封装· Anne Meixner and Laura Peters·· 5 小时前AI 评分38

Shift Left 让晶圆厂制造数据管理更复杂

Shift Left Complicates Fab Data Management

AI 导读

半导体晶圆厂借助设备传感器实时数据和 AI/ML 推理模型,将良率异常检测前移到工艺模块内部,以加速良率学习并改进先进过程控制(APC)。设备传感器参数从约每秒一次提升到每几十毫秒采样一次,数据种类较十年前至少增加 10 倍,部分晶圆厂每年产生数百 PB 数据。如何在 sub-second 处理的同时保留数据上下文与连通性,成为 shift left 带来的主要数据管理难题。

正文 · 原文

Key Takeaways:

  • “Shift left” improves yield by moving detection earlier. Fabs are using real-time equipment sensor data and AI/ML to identify process excursions inside the process module and improve advanced process control.
  • Higher sampling rates, hundreds of sensor parameters, and sub-second analysis are increasing the volume, velocity, and variety of manufacturing data.
  • Data context and connectivity remain important. However, module-level optimization affects both.

Semiconductor fabs are leveraging AI/ML models and raw equipment data to speed up yield learning, but the volume of data they are collecting — in some cases, hundreds of petabytes a year — is forcing some tough decisions about optimization versus context.

The amount of data they are wrestling with has exploded due to increased sampling from inspection and metrology tools, which is used to detect yield issues. This includes data that is streaming from equipment sensors, which is a relatively new contributor to the data volume. AI/ML inference models use this sensor data for in-situ analysis, which enables more responsive advanced process control (APC) and quicker identification of changes that could adversely impact downstream yield.

“Over the last five years, the equipment data segment has grown exponentially,” said Ashish Gupta, senior data engineering manager for automation at Intel Foundry. “This explosion is primarily the result of fabs engaging in shift-left efforts into in-situ, real-time process control.”

Others agree that the increased use of equipment data has become a source of exponential data growth.

“Equipment data has expanded along three dimensions,” said Jae Yong Park, vice president of enterprise software business at Onto Innovation. “Suppliers have added hundreds of parameters over the past few years, sampling rates have increased from about once per second to every tens of milliseconds, and engineers are increasingly working with raw trace data instead of summary data. Together, these changes have dramatically increased the volume of streaming data.”

More equipment data increases the data volume, velocity, and variety. Real-time usage of data within a process module now requires sub-second processing. Hundreds of equipment sensors have increased data variety by at least 10× from a decade ago.

“There’s been a firehose of data that’s been coming out of pretty much everywhere,” said Kartik Venkataraman, senior director of product management at Synopsys. “Even what fabs have, in terms of access to data from all their equipment suppliers, is a small fraction of the total generated by that equipment. That’s partly because OEMs want to keep some of that information to themselves. That’s their IP. A lot of times, though, it’s because the fab literally cannot handle it. If you take something like an RF plasma source in an etch chamber or a CVD chamber, that’s operating at a megahertz level. But if it’s pulsing — and pulsing is such a critical parameter — it’s operating at kilohertz. Every pulse has a waveform, which has a shape to it. It has a certain transient response and a certain droop. It has characteristics and features in there, some of which may be actionable and tell you something about your process, but you can’t possibly monitor every pulse. You’re talking about megabytes per second of data.”

For advanced device nodes, the growth in real-time/in-situ data analytics improves yield learning and advanced process control. Typically, engineers wait days, weeks, or months for a yield excursion signal from subsequent screening steps, such as electrical test. With AI/ML models, fabs can often shift identification and decisions to within the process module itself. Decisions can adjust process module settings or abort a process mid-operation. This shift requires continuous monitoring of high-resolution data on a grand scale.

Fabs shift more left
Semiconductor test has been shifting more test content left in the flow for decades.  The goal was to increase yield at final or system-level test by adding test content to wafer-level test. More recently, due to the proliferation of chiplets, an additional test step has been added for singulated die. Earlier screening increases confidence that known good die (KGD) are being shipped. This, in turn, increases the final product’s yield and quality.

Similarly, wafer fabs are shifting equipment data left in the flow, using process module data to detect yield excursions and to improve advanced process control algorithms. By relying on equipment sensor data and purpose-built virtual metrology models, process modules can quickly self-identify adverse, yield-impacting shifts within the tool. As a result, process engineers don’t need to wait for downstream screening to flag yield issues. AI/ML inference model results can identify causes of yield issues in near real-time.

“Shifting left occurs within each macro phase (Design => Fab => Sort => Advanced Packaging => Final Test),” said Intel Foundry’s Gupta. “But moving detection upstream inside the fab module, inside the design cycle, or inside the test cell is where the data engineering challenge intensifies. Instead of waiting for summary bin results or post-step inspections, in-phase shift left requires continuous, high-resolution surveillance, in-situ/in-chamber analytics, anomaly detection, virtual metrology, adaptive test, etc. These activities contribute to a 4Vs (volume, velocity, variety, veracity) data tsunami.”

Fundamentally, the shift to monitoring within a process module adds three steps. AI/ML inference models access real-time equipment sensor data. These models can flag questionable wafer processing that requires either an adjustment of equipment parameters or an abort command. For the latter, a root-cause analysis will indicate subsequent corrective actions, such as maintenance.


Fig. 1: Shift left identifies yield excursions earlier and better supports advanced process control. Source: Intel Foundry

Managing this data continues to be a critical topic as fabs increasingly request more data from the equipment suppliers. To support in-situ analysis, equipment logs need to be made available to process engineers for model building and monitoring.

“Having suppliers export key relevant data in standard formats is essential these days as chipmakers have started leveraging the data by themselves beyond traditional systems (YMS, FDC, R2R),” said Prasad Bachiraju, senior director of sales and product marketing at Onto Innovation. “Additionally, they need to pass on the knowledge of what happened to a substrate (wafer/panel) to the next step in a standard format, so information is consumed and the next process is optimized.”

Others stress the growth in real-time decisions, which requires access to equipment data to support virtual metrology models.

“Metrics should be measured inline rather than after-the-fact, because changes that could negatively impact product quality can happen at any time,” said Steve Zamek, director of product management at PDF Solutions. “We now have software that can catch these faults and interdict the equipment before making any [more] scrap, reducing the so-called product jeopardy to its absolute minimum. This is especially true in cases where we have robust virtual metrology models that can accurately predict unmeasurable in-situ product quality metrics using real-time equipment parameters.”

Facilities-related data can also benefit from this shift-left strategy. Changes to water temperature, water pressure, and power fluctuations may also affect yield. “This technique should be applied to sub-fab data as well, since the operating parameters of the components directly associated with a unit of fab equipment are having an increasing impact on the quality of the material being processed by that equipment,” Zamek noted. “The industry standards and system technologies already exist to support this kind of analysis in almost real time, so the barriers to this use case are minimal.”

Feeding data to AI/ML
AI/ML inference models are supporting the drive to use equipment sensor data for more sophisticated in-situ monitoring. Compared to multivariate statistical process control (SPC), inference models actively consume and analyze a higher data volume and a greater variety of data to support chamber analytics, anomaly detection and virtual metrology.

“It is difficult to quantify the percentage of collected data now used for analytics. However, with the advent of machine learning and artificial intelligence, fabs are collecting and using raw trace data for outlier detection and root-cause analysis rather than relying only on summary-level data,” said Onto Innovation’s Bachiraju. “The motivation is to identify subtle signals that traditional statistical analytics may have missed.”

Others confirm the increased data that is feeding today’s AI/ML inference models.

“Shifting defect detection upstream within fab modules is directly driving higher inspection frequency, denser metrology sampling, and expanded site-level measurements across the wafer. Simultaneously, breakthroughs in equipment sensor fidelity and modality are fueling rich AI/ML models that power next-generation advanced process control (APC), anomaly detection and virtual metrology applications,” said Intel Foundry’s Gupta.

The benefits of shifting left range from product yield improvement to higher overall equipment effectiveness (OEE).

“AI/ML is becoming an important driver of quality improvement because it enables organizations to move from reactive problem-solving to proactive defect prevention,” said Jon Holt, senior director of product management at PDF Solutions. “By analyzing large volumes of operational, process, and equipment data, AI models can identify patterns that are often difficult for humans to detect, helping teams predict quality issues before they occur. This improves process stability, reduces variation, accelerates root-cause analysis, and ultimately increases yield and product consistency. In complex manufacturing environments, AI also supports predictive maintenance and process optimization, reducing the likelihood of quality escapes caused by equipment degradation or process drift.”

While shift left changes what happens per operation, fabs cannot neglect screening steps or the need to connect data between process modules.

Inspection, metrology and test steps remain effective screens for what escapes these new AI/ML inference models. In addition, when connected to process module data, inspection, metrology, and test steps can identify model improvements. However, AI/ML models for connecting data between fab process steps need to operate on trustworthy data and thus traceability between these steps and context for data needs to be preserved.

Balancing data trustworthiness and traceability can be challenging, in part because standards for traceability are lacking and in part because teams are tempted to optimize data management on a per-process-module basis.

“The local optimization trap results in a situation you can describe as ‘data rich but context poor,’” said Gupta. “Each process module optimizes fiercely for its specific, time-constraint challenges. While enterprise data integration, cross-module traceability, and global correlation are more difficult to prioritize (data veracity challenge).”

Connecting data across the fab
Trading off between a process module’s optimization and the need to connect data among process modules remains a challenge. Yet ironically, connecting this data is necessary to train and maintain the AI/ML models for in-situ inference within a process module. Specifically, insufficient context for data for a process module can result as a byproduct of a process optimized for its own constraints rather than the connected data across multiple fab steps. This, in turn, can negatively impact the ability to bridge data silos within a factory and across factories.

One of the big challenges here is figuring out the right starting point. “Not only are there more tools, but they’re all generating more data, and pretty rich modes of data,” said Jim Schiely, chief of staff for Calibre Semiconductor R&D at Siemens EDA. “When you look at new giant companies like Rapidus or Terafab, there’s an opportunity to start with what we know now, rather than having little models with local ontologies and local provenance of data. For each department, they could start from scratch with the idea that this is an information business. We have to build these threads of all the different data, how they join together, and how to build this huge graph of data. These graphs are exactly the kind of fuel that AI can cook with.”

For established fabs, this requires remaking the models that are at the core of chip manufacturing. “When we did lithography models 25 years ago, we had guys who were like Michelangelos of modeling,” Schiely said. “They were the only people who knew how to do it. If you had a hard data set, you could go to these guys and they could get it working. The maturity of the modeling discipline has evolved, and the tools available for that are not only in the compute hardware, but in the libraries available in Python and the models. Automating some of the neural network models wasn’t even possible in the past. When we talked about automating model fitting 25 years ago, it seemed ridiculous. And if you couldn’t keep that data current, it wasn’t worth it to generate all these graphs and keep track of all this data, especially if it sat unused.”

AI and digital twins address these kinds of issues, but with semiconductors those abstractions need to be correlated with multi-physics simulations and testing.

“Core engines still have to generate the ground truth, and this cannot be done by LLMs,” said Synopsys’ Venkataraman. “We still have to go through the progression of first making sure that the human experts are doing it right, and that it’s validated and calibrated by real data. So the calibration piece is critical. The physical models will tell you how it should behave, but you need the actual data to anchor that in reality. Once you have that, instead of just collecting the data and seeing for yourself what you get, you understand why the data is the way it is. That’s what the digital twin is supposed to tell you. It’s supposed to reveal the physics to you. That’s where the real learning comes from, and it’s what enables generalized learning.”

That’s the goal. But data-related challenges persist across the design-through-manufacturing flow, and those challenges compound when multiple processes are involved.

“We have an increasing level of complexity with advanced devices, whether it’s gate-all-around devices in logic, 3D DRAM, modular stack NAND, and more sophisticated multi-die packaging like HBM and SoCs in a system-in-package,” said Russell Dover, managing director of equipment intelligence product development at Lam Research. “So we need to bring in smart tools to enable the digital thread, using AI technology to manage more than what we can do alone. Fundamentally, this is about speed to problem resolution and faster learning cycles.”

Engineers are going beyond process module optimization to the fleet level across fabs. “Using AI, we can create a corrective model that uses feedback or feedforward from different sorts of process tools to drive a better integrated outcome for the chip manufacturer,” said Dover. “And data does not care about physical location. And so if you have a fab, for example, in Asia, and then a sister fab somewhere else in the world, you can commingle that data and treat them as one virtual fleet. Then you can go beyond what traditional tool matching is, where you’re optimizing a fleet in a particular fab and a particular process, to now a fleet of chambers across different locations, and you can drive matching and output optimization, improving Cpk and yield for that process.”

The same applies to other processes. For example, plasma processing offers a strong platform for virtual metrology because of the ever-changing conditions associated with residue on chamber walls. “The same process chemical can behave differently in a freshly cleaned chamber than in one that has processed hundreds of wafers,” said Shalini Sharma, head of semiconductor materials innovation and cross-vertical strategy at SandboxAQ. “In plasma etch, for example, polymer residue on chamber walls can gradually change the gas-phase chemistry inside the tool and shift etch rate as the chamber ages. Engineers characterize these interactions using designed experiments, chamber telemetry, and witness wafers, then manage them through cleaning, seasoning, and run-to-run control. Machine learning can identify patterns across RF impedance, gas flow, and other sensor data that would be difficult for a person to detect, and virtual metrology can estimate results such as film thickness or etch depth before conventional measurements are available.”

Conclusion
Shifting detection and process control earlier in the manufacturing flow is accelerating yield learning, but it’s also putting a strain on the data infrastructure inside the fab. As manufacturers optimize AI/ML models within different process steps, the challenge will be preserving the context around data and the traceability needed for fab analytics and process optimization.

— Ed Sperling contributed to this article.

Related Articles
How Robotics And Intelligent Equipment Control Drive Next-Gen Fab Productivity
Behind the scenes, advanced equipment control and robotics are enabling higher levels of repeatability, tool uptime, and yield.

Secure Data Sharing Becoming Critical For Chip Manufacturing
Driven by a plethora of benefits, data sharing is gradually becoming a “must have” for advanced device nodes and multi-die assemblies.

Digital Twins: The Cloud’s The Limit
Chip industry leverages massive compute resources and AI for virtual sandboxes.

New Frontiers In Fault Detection And Classification
AI is enabling timelier and more accurate data, but AI-based command and control has yet to appear.

来源:Semiconductor Engineering 芯片与封装 · semiengineering.com