如何弥补先进芯片的“可见性鸿沟”:检测、测试与数据关联
Closing The Visibility Gap
先进芯片的每一种测试、检测与监测手段都存在物理、空间或时间上的盲区,没有任何单一技术能提供完整视图。Nordson、ZEISS、Teradyne 等厂商指出,焊点空洞、head-in-pillow 等缺陷可通过电测却导致早期失效,而 X-ray 在堆叠芯片与 TSV 检测上仍有精度、吞吐与分辨率之间的取舍。
Key Takeaways:
- No single test, inspection, model, or monitoring technique provides a complete view of an advanced device. Each has physical, spatial, temporal, or contextual blind spots.
- Important failure evidence often already exists, but it may be scattered across inspection images, test data, manufacturing history, simulation, or field measurements that were never connected.
- Better coverage increasingly depends on correlation between those views, so one measurement can narrow the uncertainty left by another rather than trying to make every tool see everything.
Advanced chips now come with an extraordinary amount of evidence about them. They’re modeled before they exist, measured during fabrication, inspected during assembly, electrically tested at the wafer and package levels, characterized under thermal and power loads, and increasingly monitored from inside the silicon after deployment. In some ways, engineers can see more of a device than ever before.
The problem is that each of those views is partial. A measurement can be accurate in every respect and still be pointed at the wrong part of the system.
“For X-ray and test technologies, we focus on the interconnect between the chips, and chips and the outside world,” said Yipeng Zhou, product line manager for Test Products at Nordson Test & Inspection. “Voids, cracks in the solder, and head-in-pillow are a few predominant defect types that pass electrical test due to the presence of physical contact. But these defects usually lead to premature failure. It’s like putting a screw in without tightening it.”
This gets at the broader visibility issue. The electrical test isn’t necessarily wrong. There’s a conductive path, and under the conditions of that test, the device may work exactly as expected. The harder question is whether the physical structure carrying that current is sound enough to keep doing so over the long term, and that’s where the visibility gap increasingly lives.
Every tool has a blind spot
As interconnects shrink, it becomes harder to catch defects before they escape. A defect that once occupied enough area to stand out might now be buried among dense stacks of bumps, wires, vias, and surrounding materials. Even when the right inspection technology exists, there’s usually a tradeoff between what engineers would like to see and what they can practically examine in production.
“For testing, there is a gap in pull testing for copper pillars and RDL integrity measurement due to the high precision requirement for the system,” said Zhou. “For X-ray, there are still gaps in inspecting overlays of wires and solder bumps from stacked dies and extremely small features in TSVs. In addition, repeatable and quantifiable measurement is a challenge for X-ray. Users also have to make choices between power, throughput, and resolution.”
Some of those limitations are simply physical. Higher resolution takes more time, greater penetration requires more power, and higher throughput means less attention for any one structure. The problem quickly becomes more complicated than finding a bad feature, because engineers first have to figure out what’s worth looking at most closely.
“It’s really not practical to just do a brute-force high-resolution X-ray scan, because high resolution means small pixels, means it’s slow, means you’ve got gigabytes worth of data, or terabytes worth of data in some cases,” said Thomas Rodgers, senior director of market strategy and head of business sector electronics at ZEISS Microscopy. “You need to have some heuristics to first know where you’re going to look.”
Failure analysis becomes less about finding one perfect tool and more about progressively narrowing uncertainty. Electrical test identifies the neighborhood, and acoustic or optical inspection might narrow it further. X-ray can look inside buried structures without immediately destroying the sample, and electron microscopy, which usually requires cross-sectioning, can then resolve the physical mechanism once engineers know where it makes sense to focus.
The same logic applies at the wafer level. A broad inspection may not tell engineers exactly what’s gone wrong, but it can show that one area, wafer, or lot looks different enough to justify a closer look. Wafer-scale macro inspection plays that screening role alongside the higher-resolution tools.
“If you’re struggling with a problem that comes with a new substrate or a more advanced node, you may need X-ray, IR, or another technique to get at that specific defect,” said Mike LaTorraca, vice president of marketing at Microtronic. “But a tool like ours can still act as a sanity check on everything. It’s your eye in the sky. It sees everything from a higher level, but you don’t stop there. It coexists with these other micro tools.”
The more useful question in production isn’t which tool can see the smallest feature. It’s whether each view reduces enough uncertainty for engineers to know what to look at next.
The missing information may already exist
Not every blind spot comes from something engineers failed to measure. Sometimes the evidence is already there, buried in a data stream that didn’t look important until something else happened later.
“Test escapes aren’t always caused because we have missing measurements,” said Eli Roth, smart manufacturing product manager at Teradyne. “The signal may already be somewhere in the wafer data.”
The signal might also be sitting in equipment telemetry, manufacturing history, final-test results, or field data. The harder part is that those records often live in different tools, different engineering groups, or completely different points in the device’s life. A number captured early in the flow has no way to announce that it’ll become important six process steps later, which is why retained history can matter so much.
A wafer that eventually breaks during a thermal process, for example, may have looked perfectly fine earlier in the flow. But if engineers have kept the inspection images, they can work backward and sometimes find that the final crack began as a much smaller deformation at the wafer edge.
“One of the keys is just having the historical images,” said Errol Akomer, applications director at Microtronic. “You may get breakage at a thermal process close to the end of line, and you can trace back with the images and look at that point of origin.”
Once the failure is known, an earlier feature that seemed inconsequential can become the most useful clue in the record. Negative evidence helps, too, and working backward through the images can be just as revealing as watching a defect develop forward.
“It’s so helpful to know where it’s not, because if you start going backward, you can blow up a certain area and keep looking at it,” said Reiner Fenske, president at Microtronic. “At some point, that little line or mark goes away. If it’s there in one image and not the previous one, then you know the problem had to start somewhere in between.”
Test data can hide clues in much the same way, especially when the result looks like good news. Engineers are naturally more likely to investigate falling yield or longer test times than a number moving in the opposite direction.
“Most commonly, anomalies are overlooked because the associated trend is normally a positive impact, rather than a negative quality or test-time impact,” said Brent Bullock, test technology director at Advantest. “If the anomaly produces lower test time or higher yield, it is more likely to be overlooked.”
Higher yield may really mean the process improved, and shorter test time may mean an optimization worked. But if no one understands why the number moved, the change still holds information that hasn’t been explained.
Sheer data volume is a little less reassuring than it sounds for the same reason. Modern test programs, inspection systems, equipment logs, embedded sensors, and manufacturing databases can produce an enormous record of what’s happening with a device, but the harder problem is figuring out which pieces of that record belong to the same story.
The model can’t show what it doesn’t know
Models are supposed to fill in some of the places engineers can’t directly see. Statistical models built on test and manufacturing data flag anomalies across huge populations of parts, while physics-based simulation estimates temperatures inside inaccessible structures, calculates voltage drop across power networks, and predicts stress through complicated material stacks. Both share the same catch, which is that a model only knows the world it has been shown.
“A model could be statistically correct, but it could still be operationally incorrect,” said Teradyne’s Roth. “The issue isn’t that the model fails itself. It’s that the model’s missing context.”
There’s another way the model can be misled too. The measurement path itself can add enough variation to obscure the thing engineers are trying to find.
“If the socket variability is making more random failure noise than any other kind of signature, you’re not going to be able to see them,” said Jack Lewis, director of applications and product management at Modus Test. “They’re hidden underneath the noise. If your raw data is noisy, it doesn’t matter. You have to filter out the noise to see the real data and build the model around it.”
Yield can look stable, control charts can look normal, and nothing may be tripping the anomaly detectors. But if the model has never seen a particular failure mechanism, or if the relevant clue lives in an earlier insertion or somewhere else in the manufacturing history, it can still miss what matters.
The missing context may not surface until much later, when something else finally goes wrong. By then, engineers may be looking back at perfectly valid earlier measurements and realizing they were only seeing pieces of the same problem.
Sometimes the missing variable is even closer to the measurement itself. Test data usually assumes the temporary interconnect is doing its job, but socket resistance can change with wear, contamination, and repeated insertions in ways that are difficult to separate from what the device is doing.
“The socket is often forgotten,” said Lewis. “There’s always this assumption that the interconnect is good, and that’s just not the case. Once you get past opens and shorts and it starts to become more contact-resistance-dependent, it’s much more difficult to see what’s going on.”
Simulation runs into a related problem, where the physics can be right while the representation is still too coarse.
“Regarding simulation accuracy against measurement, the industry never stops doing that correlation to validate the tool,” said Lang Lin, principal product manager at Synopsys. “If an EDA tool provides thermal simulation, the customer will challenge you. ‘You’ve got to provide some correlation data with my silicon so I can trust the simulation.’ That’s been done, and the correlation quality keeps improving.”
A package-level thermal model doesn’t need to predict every logic gate to be useful. Once an embedded sensor starts measuring behavior close to the device layer, though, the comparison gets much more demanding. A model that treats a die as one big thermal block may simply be looking at the problem from too far away.
“You cannot just do the single-die simulation or single-package simulation or single PCB,” said Lin. “You have to assemble them, make them exactly the same structure and positions, stacking up with the measurement device.”
New materials add another wrinkle, because values that work in one model may not transfer cleanly into the next package.
“Advanced packaging makes some of the underlying physical properties, such as thermal conductivity or Young’s modulus, much harder to characterize,” added Lin.
The process keeps looping back on itself. Measurements expose where the model is weak, the model gets adjusted, and the next round of measurements tests those assumptions again. The better engineers get at seeing the real device, the more clearly they can see where the model is still guessing.
Averages can hide the problem
Even with the right measurement, where and how engineers take it can change what they see. Temperature is a good example. A package can stay comfortably inside its average thermal limit while a much smaller region runs hot enough to create reliability problems.
“Average temperature often hides the conditions that drive reliability concerns,” said Andras Vass-Varnai, 3D-IC reliability solution architect at Siemens EDA. “A package may satisfy an average temperature target while individual regions operate substantially hotter than the package mean.”
The problem gets harder in 2.5D and 3D packages, where one hot region can affect neighboring dies, and workloads can move the hotspot around. A temperature reading that looks perfectly reasonable at one location may say very little about what’s happening a few millimeters away.
On-chip monitors can get much closer to those local conditions, but only if they’re placed where the useful information actually is.
“Some of the older technologies use the classical thermal sensor, which is typically placed at the edges of the chip because it requires an analog power supply and it’s huge,” said Nir Sever, senior director of business development at proteanTecs. “Usually you place it on the edges and try to guess what the temperature is in the center of the chip, but with the chips of today that’s no longer sufficient. You have to start probing the temperature in the hotspots that can be in many locations in the chip.”
Even dense sensor coverage doesn’t eliminate every blind spot. Sever pointed to SRAM as one of the places where internal behavior remains difficult to observe directly. “We can monitor the inputs to the memory. We can monitor the outputs from the memory. We don’t know what’s happening inside the memory.”
An error during a write, for example, may not become visible until the data is read later. The signals going in and coming out are observable, but what happened between them may stay hidden. Every technique draws a boundary like this somewhere, whether it’s an inspection that can’t scan everything at full resolution, a tester that can’t exercise every condition, or a monitor that can’t be placed inside every structure.
Connecting the dots
The growing number of measurement technologies is both the solution and part of the problem. Engineers now have electrical test, X-ray, optical and acoustic inspection, equipment data, simulation, embedded sensors, manufacturing history, and field results. The missing piece increasingly isn’t another source of data — it’s knowing how one observation relates to another.
“I don’t think right now the answer is we need more data,” said Roth. “The larger challenge is we need to connect the data and give it context, and be able to act on it quickly. We already generate enormous amounts of data.”
Failures don’t stay neatly inside one engineering discipline. A structural defect may first appear as an electrical problem, or a field failure may suddenly make an earlier test result look a lot more interesting, even though neither record seemed unusual on its own.
“How can we connect wafer sort and final test, system-level test, manufacturing history, and field performance? All that stuff is where patterns are going to emerge that are invisible if you’re looking at them independently,” Roth said. “I don’t think the answer is going to be building bigger, giant, smarter, more capable models. What we need to be doing now is creating connected data, and connected ecosystems that enable closed feedback loops.”
More coverage doesn’t necessarily mean making every test longer or putting a monitor everywhere. Test and monitoring are useful precisely because they see different parts of the problem.
“You cannot avoid test by having monitoring, for the same reason you can’t say, ‘I have the greatest test, so I don’t need monitoring,’” said proteanTecs’ Sever. “They tend to look into different things. Test is about quality — how your chip behaves versus its intended behavior and performance at time zero. For that you absolutely need test. But even if you have the best test in the world, then you move from quality to reliability. Reliability means that the quality you had at time zero remains through the useful life of the chip. For that you need monitoring.”
The handoffs between those views are improving, too, as the correlation between modeled and measured behavior keeps tightening.
“We are closing the gap,” said Synopsys’ Lin. “It’s already much better than yesterday.”
Conclusion
There probably isn’t a point where all of these blind spots disappear. Every technique makes choices about what to measure, where to measure it, at what resolution, and for how long. As chips become more heterogeneous and interactions get harder to separate, some of the most important information may keep showing up outside the place engineers first thought to look.
What’s changing is the expectation that any one measurement has to tell the whole story. An electrical result can point toward a physical defect, an inspection image can send engineers backward through the process history, a sensor can expose a hotspot that an average temperature concealed, and a field failure can give new meaning to data collected months earlier. The value often lies in the connection between those observations, rather than in any one of them alone.
Closing the visibility gap looks different as a result. The goal isn’t one test, model, sensor, or inspection system that somehow sees everything. It’s getting better at recognizing what each view can’t see and making sure another view can take the investigation the rest of the way.
Related Articles
How Long Does A Measurement Remain Valid?
Measurements can be misleading after the chip, the test path, or the assumptions behind the result change.
Smart Test Collides With The Data Chain
Increasing complexity is limiting the ability of machine learning models to effectively utilize test data.
When The Test Cell Lies
As margins shrink and dies move into expensive packages, separating device failures from test-cell artifacts has become a first-order economic problem.
来源:Semiconductor Engineering 芯片与封装 · semiengineering.com