Introduction

Clean-energy industrial policy has transformed renewable-energy law from a permitting and subsidy field into a national-security field. Battery cells, cathode and anode materials, grid devices, charging networks, hydrogen systems, nuclear inputs, and critical-mineral recycling are no longer treated only as climate infrastructure. They also function as supply-chain chokepoints, sources of operational data, and possible points of leverage over U.S. industrial capacity. Foreign investment therefore sits at the center of a legal tension. The same capital and technical expertise that can accelerate decarbonization may create foreign-control, technology-transfer, data, cybersecurity, or resilience concerns. The Committee on Foreign Investment in the United States, or CFIUS, is the principal U.S. mechanism for reviewing whether covered foreign investment affects national security [1], [2].

FIRRMA enlarged CFIUS jurisdiction beyond traditional control acquisitions to include certain non-controlling investments in TID U.S. businesses, with TID referring to critical technology, covered investment critical infrastructure, and sensitive personal data [2], [3]. The 2020 implementing regulations made this jurisdiction operational and connected mandatory declarations to critical technology and substantial foreign-government interests [3]. Executive Order 14083 then directed CFIUS to consider evolving risks, including supply-chain resilience, critical minerals, cybersecurity, sensitive data, and cumulative or sectoral investment patterns [5], [6], [13]. These legal materials make renewable-energy supply chains an appropriate field for AI-assisted screening: the inquiry is text-rich, factor-based, and highly dependent on factual classification.

Renewable-energy projects are especially suitable for this type of screening because they combine several CFIUS-relevant dimensions in a single transaction record. A battery plant can involve critical minerals, cell manufacturing know-how, production software, energy-storage integration, foreign supply agreements, and a large local employment footprint. An EV investment can involve vehicle assembly, charging networks, telematics, software updates, battery sourcing, and joint-venture governance. A grid or storage project can involve operational technology and cybersecurity risks even when the immediate investment does not involve a defense contractor. The legal issue is therefore multi-factorial: the same project may be relevant to critical technology, critical infrastructure, supply-chain resilience, sensitive data, export controls, and real-estate proximity. A matrix that treats these dimensions separately is more transparent than a single unexplained risk label.

The energy transition also creates a timing problem. Investors need to decide whether to file a declaration, prepare a full notice, adjust governance rights, change supply terms, or design mitigation before construction schedules and subsidy deadlines become fixed. CFIUS review is therefore not only a post-signing compliance issue; it is an ex ante planning issue. The practical value of legal AI is to convert scattered public legal materials into a repeatable checklist that counsel can apply at the term-sheet, site-selection, and supply-contract stages [28]-[30]. This paper evaluates that practical claim with data rather than with a conceptual example.

The rule-of-law problem is not whether CFIUS should consider clean-energy supply-chain risks. The problem is how investors, counsel, public agencies, and affected communities can anticipate review intensity when national-security standards are necessarily flexible and much of the government record remains confidential. Rule-of-law theory stresses generality, publicity, prospectivity, clarity, consistency, and congruence between announced rules and official action [16]-[19]. CFIUS cannot publish confidential risk intelligence, but it can be studied through public legal sources, annual reports, and project data. A transparent screening matrix can therefore improve ex ante predictability while preserving CFIUS discretion for case-specific threats.

This article makes three contributions. First, it converts public CFIUS and energy-supply-chain materials into a reproducible legal-AI taxonomy for renewable-energy investment. Second, it conducts full experimental evaluations on the specified project and legal-document datasets. The results reported below are measured by executable scripts and are not estimates. Third, it links the empirical results to rule-of-law criteria: transparency, predictability, proportionality, reason giving, and procedural safeguards. The framework does not claim to predict confidential CFIUS decisions. It predicts a screening class generated by an auditable coding protocol, which is the proper empirical target when actual classified or confidential CFIUS outcomes are unavailable.

The central research question is whether a reproducible AI-assisted legal framework can identify CFIUS-relevant risk features in renewable-energy supply-chain projects and thereby support more predictable preliminary screening. The answer is affirmative for triage, issue spotting, and legal planning, but not for final jurisdictional or national-security conclusions. The evidence shows that technology type, critical-mineral exposure, manufacturing role, private-capital status, software/data exposure, and investment scale consistently separate high-sensitivity battery and EV projects from lower-sensitivity deployment grants under the disclosed protocol. The framework also shows where public project data remain under-specified: investor nationality, beneficial ownership, actual control rights, parcel-level real-estate proximity, export-control classification, mitigation history, and non-public threat intelligence still require lawyer and agency judgment.

This framing also distinguishes legal AI from ordinary compliance automation. A compliance checklist usually asks whether a rule is triggered. A CFIUS screen must first identify whether facts are missing, whether multiple legal regimes overlap, and whether risk is better addressed through filing, governance design, supply-chain restructuring, or mitigation. The AI component in this paper therefore operates as structured legal reasoning: it converts public legal factors into observable variables, measures how those variables distribute across a project corpus, and tests whether the resulting classifications can be reproduced by models that counsel can audit [31], [33], [36].

Method

The method was designed as a reproducible legal-AI experiment, but its legal purpose is practical: to create a documented preliminary screen that a transaction team can review before deciding whether more intensive CFIUS diligence is needed. The workflow uses deterministic preprocessing, topic clustering, an auditable risk score, supervised model comparison, ablation analysis, and feature-importance explanations. Proprietary LLM calls were not used for the reported numerical results because undisclosed model weights and prompts would reduce reproducibility. The workflow is LLM-ready: the same coded features and extracted rationales can be supplied to a language model for drafting memoranda [35], [37], but the measurements in this paper come from transparent NLP and machine-learning scripts. This design follows the interpretability concerns raised in machine-learning literature for high-stakes settings [20]-[24], [26], [38], [42], [44].

Table I. Data sources and analytical roles.

Dataset Files Count Fields or content Analytical role
Project-level clean-energy table data.csv 2,064 Technology, state, investment, jobs, public/private indicators Risk-scoring and supervised classification
CFIUS annual reports CY 2020-CY 2024 report PDFs 5 Declarations, notices, investigations, mitigation, withdrawals Procedural-rate trends and legal corpus
Core CFIUS legal texts FIRRMA and Part 800 PDF 2 Jurisdiction, covered transactions, TID businesses, critical technologies Feature taxonomy
Evolving-risk and energy-supply-chain texts EO 14083, FEOC, clean-vehicle and battery PDF files 4 Supply-chain resilience, critical minerals, FEOC, clean-vehicle constraints NLP topic clustering and rationale mapping

Table I lists the empirical materials. The project-level table provided the structured observations. The legal-document corpus provided the terms used to build the risk taxonomy. CFIUS annual reports supplied procedural context on declarations, notices, investigations, mitigation, withdrawals, and presidential decisions. FIRRMA and Part 800 supplied jurisdictional concepts, including covered transactions, critical technologies, covered investment critical infrastructure, and sensitive data [2], [3]. Executive Order 14083 supplied the national-security expansion toward supply-chain resilience, critical minerals, cybersecurity, and cumulative sectoral exposure [6]. Battery and FEOC materials supplied the clean-vehicle and critical-mineral connection [12], [14], [15].

Table II. Project dataset profile after cleaning.

Measure Value
Project records 2,064
Unique states/territories 55
Technology categories after normalization 12
Manufacturing records 1,105
Private-capital records 846
Public-capital records 1,218
Total investment, USD 300,202,834,967
Private investment, USD 198,548,610,800
Public investment, USD 101,654,224,167
Reported jobs 209,268
Records with missing city 1,287

The project table contains 2,064 records after technology normalization. The only normalization that changed a category label combined 'Offshore wind' with 'Offshore Wind.' The dataset contains both public and private investment indicators. It also includes state, city, investment, jobs, technology, project, company name, and category fields. Latitude and longitude were retained in the replication file but excluded from scoring because the records do not provide parcel-level real-estate proximity facts required for a reliable Part 802 analysis [4]. Missing city values were coded as a due-diligence uncertainty rather than as proof of proximity risk.

The risk score is a statutory-feature screening score, not a government outcome label. Each record receives points for legal factors that would matter if foreign participation were present. The score is conditional in that sense: a U.S.-only project may have high national-security sensitivity but no CFIUS filing obligation unless a covered foreign person, covered investment, covered control transaction, or covered real-estate transaction is involved [2]-[4]. The score therefore answers a planning question: how intensive should preliminary CFIUS diligence be for this project profile?

Table III. Auditable CFIUS screening score components.

Risk component Operational weight Coding rationale
Technology sensitivity Batteries 26; EV 21; grid 22; nuclear 24; hydrogen 15; solar 12; other technologies 6-10 Legal sensitivity under TID, critical infrastructure, and supply-chain factors.
Critical minerals/battery exposure +15 Battery materials, lithium, graphite, cobalt, nickel, manganese, cells, cathodes, anodes, recycling, and EV supply-chain terms.
Manufacturing and supply-chain role +12 Production facilities and component manufacturing create dependency and resilience concerns.
Investment channel +12 private; +3 public Private-capital records receive the stronger CFIUS-relevance proxy; public grants receive lower sensitivity.
Investment scale +2 to +15 by sample quantile Larger projects increase operational and market dependency if foreign-controlled.
Job/operational footprint +3 to +8 by reported jobs High employment and operating footprint indicate materiality.
Software/data exposure +6 batteries; +8 EV/nuclear; +10 grid Networked grid, charging, BMS, and operational software create cybersecurity and data pathways.
Location uncertainty +3 missing city; +5 unreported state Incomplete site detail limits real-estate proximity review and is treated as due-diligence uncertainty.

Table III operationalizes the legal matrix. Technology sensitivity receives the largest baseline weight because CFIUS law focuses on critical technologies and infrastructure, while EO 14083 adds supply-chain resilience and critical minerals [3], [6]. Battery and critical-mineral exposure receives a separate weight because lithium, graphite, cobalt, nickel, manganese, recycling, cells, cathodes, and anodes connect clean-energy projects to upstream dependency. Manufacturing receives a separate weight because the legal concern is not only technology ownership but the location and control of production capacity. Private-capital status receives a larger weight than public grants because CFIUS jurisdiction is triggered by foreign investment, not by ordinary public funding. Software/data exposure captures grid, charging, BMS, sensor, and operational-data pathways that can raise cybersecurity and sensitive-data concerns [32], [48]-[50], [52].

Risk classes were assigned by fixed thresholds: Low for scores below 40, Medium for scores from 40 through 64, and High for scores of 65 or above. The thresholds were set before model training and were not tuned to maximize classifier performance. The classifiers were trained only after the labels were assigned. This sequencing prevents the model from defining the target it later predicts.

The coding protocol intentionally separates sensitivity from jurisdiction. Sensitivity asks whether the project contains attributes that public CFIUS materials identify as national-security relevant. Jurisdiction asks whether the transaction involves a foreign person, control, covered investment rights, substantial foreign-government interest, or covered real estate. The project table supports the first question and partially supports the second through its private and public investment fields. It does not provide ownership-chain data, so the experiment does not turn sensitivity into a filing obligation. This distinction protects the legal validity of the empirical design.

The score also separates positive risk evidence from data-quality uncertainty. Missing city values do not prove a proximity problem. They do show that a real-estate screen cannot be completed from the record alone. The model therefore assigns a small uncertainty weight rather than a substantive proximity weight. That choice matters for rule-of-law reasons: a transparent model should not punish a project as dangerous when the dataset merely lacks a field needed for further diligence [38], [54].

The legal text experiment extracted text from PDF files, split it into paragraphs with at least 35 words, and capped each source at 250 retained paragraphs to prevent a single long regulation from dominating the corpus. TF-IDF vectors used one- and two-word terms with a maximum of 1,500 features. K-means clustering with six clusters grouped legal paragraphs into issue families. A cosine-silhouette statistic measured separation, and top terms plus dominant sources were used to assign human-readable labels. The purpose was not to discover a perfect ontology. It was to test whether a reproducible NLP pipeline could recover the main doctrinal issue families that counsel would use in a CFIUS screening memorandum [35], [37].

The supervised experiment compared seven models: a majority-class baseline, TF-IDF Naive Bayes, TF-IDF logistic regression, structured logistic regression, a linear SVM using text and structured features, a structured random forest, and structured XGBoost. The split was stratified 70 percent training and 30 percent holdout testing with random seed 42. Evaluation metrics were accuracy, balanced accuracy, macro-F1, High-risk recall, and High-risk precision. The strongest model was then evaluated with five-fold stratified cross-validation. Ablation analysis removed feature families to measure which facts carried screening signal. Tree-gain or coefficient importance supplied interpretable rationales for the highest-performing structured models [40], [44], [46], [47].

Macro-F1 was selected as the main comparison metric because the classes are imbalanced: Low records are the largest class, while High records are legally the most consequential. A model that predicts only Low would have non-trivial accuracy but would fail the legal task. High-risk recall was therefore reported separately. In CFIUS triage, a false negative that sends a high-sensitivity project to the lowest review lane is more costly than a false positive that sends a project to additional legal review. Balanced accuracy was included to show whether models treated the three classes evenly rather than exploiting the majority class [46], [47].

Results and Discussion

The project corpus is concentrated in the technologies that most strongly connect clean-energy policy to national-security screening. Table IV shows that batteries and EVs account for USD 194.71 billion of the USD 300.20 billion total investment, or 64.9 percent of total investment in the project table. Batteries alone account for USD 140.44 billion, 368 records, and 79,907 reported jobs. EVs account for USD 54.27 billion, 207 records, and 66,965 reported jobs. The concentration is legally important because batteries and EVs combine manufacturing, critical-mineral dependency, software, charging, data, and supply-chain resilience issues. This is exactly the combination that FIRRMA, Part 800, EO 14083, and FEOC materials bring into legal focus [2], [3], [6], [14].

Table IV. Technology profile and screening outcomes.

Technology Records Total investment USD Jobs Median score High-risk share %
Batteries 368 140,441,693,175 79,907 76 97.6
Electric Vehicles 207 54,273,531,762 66,965 76 100
Other 859 40,428,225,322 0 13 0
Solar 152 21,332,250,000 43,966 44 2.6
Nuclear 47 14,382,817,389 1,905 56 17
Grid 126 10,133,514,248 0 43 0
Hydrogen 78 9,262,206,154 5,510 39 0
Offshore Wind 80 7,191,454,550 3,875 37 0
Fossil 38 1,272,019,398 0 21 0
Heat Pumps & Clean HVAC 33 975,733,361 3,246 39 0
Land-Based Wind 18 465,200,000 3,894 35 0
Hydroelectric 58 44,189,608 0 28 0

Figure 1. AI-assisted CFIUS screening workflow.

Figure 2. Project records by clean-energy technology.

Figure 3. Investment concentration by technology.

The descriptive results show a sharp difference between deployment-heavy public programs and manufacturing-heavy supply-chain projects. Public records dominate the count, with 1,218 public-capital records, but private-capital records carry most of the high-sensitivity manufacturing investment. This distinction matters for rule-of-law analysis. A purely sectoral rule that treats every clean-energy project as equally sensitive would over-screen small public grants and under-explain why large battery plants, EV manufacturing projects, and critical-mineral facilities need enhanced diligence. The matrix makes the distinction explicit.

The investment totals also show why a CFIUS screen cannot be limited to formal technology categories. Some records classified as Other contain manufacturing or industrial-facility indicators but lack battery, EV, grid, or mineral terms. Those records normally remain Low because they do not combine enough legally relevant features. Conversely, battery and EV records become High not merely because the labels are politically salient, but because the records combine large private investment, manufacturing, jobs, critical-mineral exposure, and software or data pathways. The matrix therefore supports a fact-based explanation that can be reviewed by counsel, investors, and public stakeholders.

Table V. Risk class distribution by technology.

Technology Low Medium High Total
Batteries 0 9 359 368
Electric Vehicles 0 0 207 207
Nuclear 0 39 8 47
Solar 49 99 4 152
Grid 0 126 0 126
Fossil 37 1 0 38
Heat Pumps & Clean HVAC 20 13 0 33
Hydroelectric 58 0 0 58
Land-Based Wind 13 5 0 18
Hydrogen 41 37 0 78
Offshore Wind 51 29 0 80
Other 830 29 0 859

Figure 4. Risk class composition by technology.

The risk-class distribution is consistent with the legal assumptions embedded in the screening protocol. Under the disclosed scoring rules, every EV record is classified High, and batteries are almost entirely High, with 359 of 368 records classified High. Grid projects are classified Medium because their cybersecurity and infrastructure features create sensitivity, but the dataset does not contain private-capital or mineral facts strong enough to push those records into High [48]-[53]. Nuclear projects split between Medium and High depending on scale and related features. Other, hydroelectric, fossil, wind, and many public deployment records cluster in Low because they lack the combined investment, manufacturing, critical-mineral, and software characteristics that drive the protocol. These are screening results, not external findings about CFIUS filing obligations.

The annual-report extraction adds institutional context. Table VI reports measured CFIUS procedural pressure indicators for calendar years 2020 through 2024. The highest notice volume in the extraction was 286 notices in 2022; the highest notice investigation rate was 56.6 percent in 2022; and the highest notice mitigation rate was 15.0 percent in 2023. Declaration clearance rose from 64.3 percent in 2020 to 78.4 percent in 2024, while notice withdrawal rates remained substantial. These values support the legal point that declarations can clear lower-risk matters, but notice practice remains important for complex transactions. Earlier annual reports already document the post-FIRRMA shift toward declarations and expanded review authority [7]-[9].

For renewable-energy investors, those procedural indicators have concrete planning consequences. A declaration can be efficient where the ownership chain is simple, the technology is not at the core of critical minerals or sensitive infrastructure, and mitigation is unlikely. A full notice is more appropriate when the project combines battery materials, foreign control rights, sensitive software, large-scale manufacturing, or uncertain supply-chain dependencies. The empirical model cannot choose the filing form by itself, but it provides a defensible triage record that counsel can use when documenting why a declaration, notice, or no-filing strategy was selected [10], [11].

Table VI. CFIUS annual procedural indicators extracted from the legal corpus.

Year Decl. Mandatory decl. Notices Investigation % Mitigation % Withdrawal % Pres. decisions
2020 126 34 187 47.1 8.6 15.5 1
2021 164 47 272 47.8 9.6 27.2 0
2022 154 44 286 56.6 14.3 30.8 0
2023 109 36 233 54.9 15 24.5 0
2024 116 36 209 55.5 7.7 23.4 2

Figure 5. CFIUS procedural pressure indicators.

The legal NLP experiment recovered six reproducible issue families from the legal corpus. Table VII lists top terms and dominant sources. The silhouette value of 0.087 shows that the legal corpus uses overlapping vocabulary, especially the recurring terms 'transaction,' 'national security,' 'foreign,' 'notice,' and 'critical.' The low separation is informative rather than defective: CFIUS materials are legally integrated, and the same statutory vocabulary appears in procedure, mitigation, technology, and country tables. Top-term inspection nevertheless separates FEOC battery-credit rules, covered-notice and sector tables, CFIUS mitigation and transaction process, critical technology and TID jurisdiction, manufacturing taxonomy, and country/economy declaration trends [31], [33], [34], [37].

Table VII. Legal corpus topic clusters.

Cluster Issue family Paragraphs Dominant source Top terms Silhouette
1 Tax-credit and FEOC battery rules 135 2024-09094.pdf 30d, vehicle, battery, section, credit, 25e, section 30d, taxpayer, proposed, regulations 0.087
2 Covered notices, declarations, and sectors 85 CFIUS - Annual Report to Congress CY 2022_0.pdf notices, percent, sector, transactions, accounted, cfius, subsector, table, 2021, accounting 0.087
3 CFIUS mitigation and transaction process 282 2023CFIUSAnnualReport.pdf cfius, transaction, committee, parties, notice, security, national security, mitigation, national, transactions 0.087
4 Critical technology and TID jurisdiction 360 Part-800-Final-Rule-Jan-17-2020.pdf foreign, states, united states, critical, united, security, national, section, technologies, 800 0.087
5 Business-sector and manufacturing taxonomy 87 CFIUS-Public-AnnualReporttoCongressCY2021.pdf manufacturing, 100, services, naics, 100 100, transportation, product, activities, digit, support 0.087
6 Country/economy declaration trends 60 2023CFIUSAnnualReport.pdf economy, country, country economy, transactions, 10, 11, values, islands, 13, 12 0.087

Figure 6. Legal corpus topic separation.

The supervised models were evaluated against the auditable screening classes. Table VIII shows that text-only models performed well because technology and company terms contain strong signals. TF-IDF logistic regression reached 0.931 accuracy and 0.922 macro-F1. Structured models performed better because the labels use structured legal features. Structured logistic regression reached 0.982 accuracy and 0.979 macro-F1. The structured random forest was the best holdout model, with 0.989 accuracy, 0.986 macro-F1, 0.994 High-risk recall, and 1.000 High-risk precision. XGBoost reached the same accuracy and 0.985 macro-F1. These results are coherent with the task: the models reproduce a transparent legal coding protocol, not a hidden agency decision.

Table VIII. Holdout model comparison.

Model Accuracy Balanced accuracy Macro-F1 High recall High precision
Majority baseline 0.532 0.333 0.232 0 0
TF-IDF Naive Bayes 0.913 0.922 0.901 0.983 0.977
TF-IDF Logistic 0.931 0.946 0.922 0.983 1
Structured Logistic 0.982 0.984 0.979 0.994 1
Linear SVM full 0.987 0.987 0.983 0.983 1
Random Forest structured 0.989 0.981 0.986 0.994 1
XGBoost structured 0.989 0.982 0.985 0.989 1

Figure 7. Holdout model performance by model.

Cross-validation confirmed that the holdout result was stable. Table IX reports five stratified folds for the best model. Mean accuracy was 0.991, mean macro-F1 was 0.988, and mean balanced accuracy was 0.986. The standard deviation of macro-F1 was 0.004. The narrow dispersion shows that the screening features separate the classes consistently across folds. The confusion matrix in Table X and Figure 8 shows the error pattern: all 330 Low records in the holdout set were correctly predicted; 110 of 116 Medium records were correctly predicted; 173 of 174 High records were correctly predicted. The only High-risk error was classified as Medium, not Low, so the model did not suppress a high-sensitivity project into the lowest triage category.

Table IX. Five-fold cross-validation for the best model.

Fold Accuracy Macro_F1 Balanced_accuracy
1 0.988 0.984 0.984
2 0.995 0.994 0.991
3 0.993 0.991 0.988
4 0.988 0.985 0.98
5 0.99 0.988 0.987
Mean 0.991 0.988 0.986
Std. 0.003 0.004 0.004

Table X. Holdout confusion matrix for the best model.

Class Predicted Low Predicted Medium Predicted High
Actual Low 330 0 0
Actual Medium 6 110 0
Actual High 0 1 173

Figure 8. Best-model holdout confusion matrix.

Ablation analysis identifies the source of legal signal within the disclosed coding protocol. Removing capital scale reduced macro-F1 from 0.986 to 0.928. Using only capital scale reduced macro-F1 to 0.647, which indicates that investment size alone is insufficient. Using only technology and category reached 0.891 macro-F1, which indicates that sector classification carries substantial signal but still needs scale, mineral, software, and capital-channel facts. Removing investment channel did not change performance in this dataset because technology and mineral features already identify most High-risk records. In a dataset with actual investor nationality and control rights, investment-channel variables would be expected to matter more for jurisdiction.

Table XI. Feature ablation results.

Feature set Accuracy Macro-F1 High-risk recall
Full structured 0.989 0.986 0.994
No technology/category 0.984 0.98 0.994
No capital scale 0.937 0.928 0.983
No investment channel 0.989 0.986 0.994
Technology/category only 0.902 0.891 0.983
Capital scale only 0.721 0.647 0.822

Table XII. Top feature-importance results.

Feature Importance
critical_mineral_exposure 0.162
software_data_exposure 0.126
tech_norm_Solar 0.104
tech_norm_Grid 0.102
private_capital_flag 0.067
tech_norm_Other 0.046
log_priinvest 0.038
priinvest 0.033
tech_norm_Hydroelectric 0.031
category_All 0.029
manufacturing_flag 0.027
state_Unreported 0.027
state_unreported_flag 0.023
tech_norm_Hydrogen 0.022
tech_norm_Nuclear 0.02

Figure 9. Most influential screening features.

Feature-importance results are legally coherent. Critical-mineral exposure is the strongest feature, followed by software/data exposure, solar and grid technology indicators, private-capital flag, other-technology classification, and private-investment scale. Solar appears with a high importance value because it helps distinguish Medium and Low records from battery and EV records; in other words, importance can reflect negative separation as well as positive risk. Grid appears because it separates infrastructure and cybersecurity sensitivity from ordinary public clean-energy records. The importance table therefore functions as a concise rationale generator: when a project is classified High, the analyst can identify whether the driver is mineral dependency, software/data exposure, private capital, manufacturing, scale, or technology type [44], [45].

The high performance of the structured models should be read together with the ablation results. The models do not reveal a hidden natural law of CFIUS outcomes. They show that the disclosed legal coding protocol is internally consistent and machine-reproducible. That is exactly the result a rule-of-law screening tool should deliver. If two analysts apply the same public factors to the same project table, they should reach the same triage result or be able to identify which input caused the difference. The model's value is reproducibility, auditability, and error detection, not autonomous legal authority.

The rule-of-law implication is that AI assistance improves legal planning when it is transparent, bounded, and contestable. The score does not replace the CFIUS legal standard. It translates public legal factors into a disclosed screening protocol [25], [27], [33], [39]. Investors can challenge inputs, adjust assumptions, and document why a project should be treated as lower or higher risk. Agencies can also benefit because parties that self-screen with a transparent matrix can file clearer declarations or notices, identify mitigation topics earlier, and avoid treating national security as an entirely opaque label.

The framework also supports proportionality. High-risk battery and EV manufacturing projects should receive early CFIUS diligence, ownership-chain review, supply-chain mapping, cybersecurity analysis, and mitigation planning. Medium-risk grid and nuclear-adjacent records should receive targeted infrastructure and software review. Low-risk public deployment records should not be overburdened absent foreign-control, proximity, data, or critical-technology facts. This tiered approach aligns CFIUS screening with the rule-of-law value that legal burdens should be related to legally relevant risk rather than imposed on an entire policy sector.

The same structure improves reason giving. A national-security review system can remain confidential while still giving parties clearer categories of concern. A screening memorandum based on this matrix can state, for example, that the concern is not 'clean energy' in general but battery-material dependency, foreign governance rights, grid operational data, or the inability to complete a proximity screen [28]-[30], [34]. Narrow reasons are easier to contest and easier to mitigate. They also reduce the danger that national security becomes a conclusory phrase detached from statutory factors.

The results also show how AI can be used without undermining legal accountability. The legal judgment remains with counsel and the agency. The model performs repeatable classification, records the features that matter, and exposes the assumptions that require factual verification. This division of labor is consistent with the rule-of-law demand that consequential decisions be explainable by legal reasons rather than by opaque computation [27], [39], [41], [43].

Several legal-reform implications follow from the measured results. Public guidance could define examples of battery, critical-mineral, grid-software, and EV supply-chain transactions that normally require enhanced diligence. Clean-energy subsidy programs could collect standardized ownership, governance, and site-proximity fields so that national-security review does not begin with missing data. CFIUS practice could also encourage parties to submit structured risk matrices with declarations and notices. These reforms would not disclose classified information. They would improve the public-facing grammar through which parties organize facts and through which the government communicates the categories of concern that matter.

Practical Implications for Transaction Counsel and Policymakers

For transaction counsel, the framework is most useful at the front end of a deal, before filing strategy, governance design, supply-chain diligence, and mitigation planning become locked into the transaction timetable. A battery plant, EV manufacturing investment, grid-software acquisition, or critical-mineral recycling project can be screened for the features that public CFIUS materials identify as relevant: foreign investment channel, manufacturing role, critical minerals, technology sensitivity, software/data exposure, investment scale, jobs, and location uncertainty. The resulting score should not decide whether to file [28], [34]. It should identify which factual workstreams need immediate review.

The framework also helps counsel separate sensitivity from jurisdiction. A project may be nationally sensitive because of minerals, grid operations, or software, but it does not become a CFIUS-covered transaction unless the transaction facts involve a covered foreign person, control rights, covered investment rights, substantial foreign-government interests, or covered real estate. That distinction is important in practice because it prevents the model from turning sector salience into a legal conclusion. The screen should therefore be paired with ownership-chain diligence, governance-rights review, export-control analysis, real-estate proximity review, and supply-contract analysis [34].

For policymakers, the results show that clean-energy national-security review would benefit from more standardized public data. Project datasets and subsidy programs could collect ownership, governance, site-proximity, technology-control, and supply-chain fields in a consistent format. That would not require disclosure of classified threat intelligence. It would improve the public-facing grammar through which parties organize relevant facts and through which government agencies communicate the categories of concern that matter.

For investors and communities, the rule-of-law value is proportionality. A transparent matrix can explain why a battery-materials facility or EV manufacturing project should receive enhanced diligence while a public deployment grant may not need the same burden absent foreign-control, proximity, data, or critical-technology facts. The objective is not to narrow CFIUS discretion mechanically, but to make preliminary self-assessment more public, contestable, and tied to legally relevant factors.

Limitations

The empirical target is a screening class generated by an auditable legal coding protocol. It is not an actual CFIUS decision label. Actual CFIUS outcomes are influenced by classified intelligence, non-public party submissions, mitigation negotiations, ownership rights, governance covenants, export-control facts, buyer identity, and agency threat assessments. Those facts are not present in the project table. The correct interpretation is therefore limited but useful: the model reliably reproduces the study's screening protocol; it does not predict confidential government action [27], [29], [33].

This limitation also explains why the article does not report a conventional accuracy score against CFIUS approvals, withdrawals, mitigation agreements, or presidential decisions at the individual-project level. Such labels are not present in the dataset and are not publicly available for the project table. Creating them would have produced a false empirical claim. The study instead evaluates whether public legal factors can be translated into a consistent screening target and whether models can reproduce that target with measured performance.

The project table does not contain investor nationality, beneficial ownership, voting rights, board rights, veto rights, supply agreements, technology-control agreements, or foreign-government interests. The risk score uses private-capital status as a proxy for investment-channel relevance, but it does not determine whether a transaction is a covered control transaction, covered investment, excepted-investor transaction, or covered real-estate transaction under Part 800 or Part 802 [3], [4]. A complete legal memorandum must add ownership-chain diligence.

The dataset structure also means that the model cannot assess country-specific threat, sanctions exposure, export-control classification, or foreign-government ownership. Those are not minor details in CFIUS practice. They are often decisive. The screening score should therefore be used at the beginning of diligence, not at the end. A High score identifies projects for deeper legal review; a Low score does not eliminate review when independent facts show foreign control, military proximity, sensitive data, or critical technology.

The dataset also lacks parcel-level locations, facility boundaries, distance to military installations, ports, sensitive government facilities, or other proximity facts. Location uncertainty is coded only as due-diligence uncertainty. It is not coded as substantive real-estate risk. This limitation protects logical coherence because CFIUS real-estate analysis requires site-specific facts that cannot be inferred from state names or incomplete city fields.

The NLP clustering experiment uses public legal documents and annual reports. It does not use confidential CFIUS filings, mitigation agreements, or classified threat materials. Its six clusters are reproducible issue families, not an official CFIUS taxonomy. The low silhouette score reflects overlapping statutory vocabulary in the corpus. It does not invalidate the experiment because the clustering output is used for legal issue mapping, not for high-stakes project classification [31], [37].

Finally, the high model scores follow from the transparency of the coding protocol. In a confidential-outcome prediction task, performance would likely be lower because actual decisions depend on facts not in the table. This limitation is a strength for the article's rule-of-law purpose: the framework favors interpretable screening over opaque prediction. It identifies the facts that investors should investigate before making claims about CFIUS filing strategy.

Conclusion

Renewable-energy supply chains now sit at the intersection of climate policy, industrial policy, foreign investment, and national security. CFIUS law supplies the principal institutional mechanism for reviewing covered foreign investment, but its case-by-case and partly confidential process creates predictability challenges for investors, counsel, public agencies, and communities. This paper developed and evaluated an AI-assisted screening framework that translates public legal sources and project facts into reproducible preliminary risk classes.

The experimental results are consistent with the paper's limited empirical target. The project corpus contains 2,064 records, USD 300.20 billion in total investment, and 209,268 reported jobs. The screening matrix classified 578 records as High risk, 387 as Medium, and 1,099 as Low. Batteries and EVs carry the most concentrated sensitivity because they combine private investment, manufacturing, critical-mineral exposure, software/data pathways, and large operational footprints. The best model reached 0.989 holdout accuracy and 0.986 macro-F1, with cross-validated macro-F1 of 0.988. Ablation and feature-importance results show that critical-mineral exposure, software/data exposure, technology type, private capital, and investment scale are the principal screening drivers under the disclosed coding protocol.

The doctrinal lesson is equally concrete. National-security review becomes more legitimate when the reasons for review can be organized around public factors: control, technology, infrastructure, data, supply-chain dependency, minerals, cybersecurity, and proximity. AI assistance can help organize those factors, but it must remain subordinate to law. The method used here gives each project a contestable score, not an unreviewable conclusion [38], [40], [42], [44], [54].

The framework can be piloted in compliance practice. A transaction team can run project facts through the scoring script, review the feature drivers, verify ownership and control rights, add site-specific proximity information, analyze export-control and supply-chain facts, and then decide whether a declaration, notice, mitigation plan, or no-filing record is appropriate. The same workflow can help policymakers identify which clean-energy subsectors need clearer guidance and which public-grant programs should collect better ownership, governance, and location information.

The legal conclusion is that AI should not replace CFIUS, lawyers, or statutory judgment. It should structure the questions that precede them. A transparent screening matrix gives foreign investors a clearer basis for self-assessment, gives counsel a reproducible triage method, and gives policymakers a way to discuss national-security review without relying solely on opaque discretion. The rule-of-law payoff is not mechanical prediction. It is a more public, contestable, and proportionate pathway for aligning clean-energy investment with national-security safeguards.

References

[1] U.S. Congress, Foreign Investment and National Security Act of 2007, Pub. L. No. 110-49, 121 Stat. 246, 2007.

[2] U.S. Congress, Foreign Investment Risk Review Modernization Act of 2018, Pub. L. No. 115-232, div. A, title XVII, 132 Stat. 2174, 2018.

[3] U.S. Department of the Treasury, "Provisions Pertaining to Certain Investments in the United States by Foreign Persons," Federal Register, vol. 85, no. 12, pp. 3112-3202, Jan. 17, 2020.

[4] U.S. Department of the Treasury, "Provisions Pertaining to Certain Transactions by Foreign Persons Involving Real Estate in the United States," Federal Register, vol. 85, no. 12, pp. 3158-3197, Jan. 17, 2020.

[5] The White House, "Executive Order 14017 of February 24, 2021: America's Supply Chains," Federal Register, vol. 86, no. 38, pp. 11849-11854, Mar. 1, 2021.

[6] The White House, "Executive Order 14083 of September 15, 2022: Ensuring Robust Consideration of Evolving National Security Risks by the Committee on Foreign Investment in the United States," Federal Register, vol. 87, no. 181, pp. 57369-57374, Sep. 20, 2022.

[7] U.S. Department of the Treasury, "CFIUS Annual Report to Congress, Calendar Year 2020," Washington, DC, 2021.

[8] U.S. Department of the Treasury, "CFIUS Annual Report to Congress, Calendar Year 2021," Washington, DC, 2022.

[9] U.S. Department of the Treasury, "CFIUS Annual Report to Congress, Calendar Year 2022," Washington, DC, 2023.

[10] U.S. Department of the Treasury, "Guidance Concerning the National Security Review Conducted by the Committee on Foreign Investment in the United States," Federal Register, vol. 73, no. 236, pp. 74567-74572, Dec. 8, 2008.

[11] U.S. Department of the Treasury, "CFIUS Enforcement and Penalty Guidelines," Washington, DC, 2022.

[12] U.S. Department of Energy, Federal Consortium for Advanced Batteries, "National Blueprint for Lithium Batteries," Washington, DC, 2021.

[13] The White House, "Building Resilient Supply Chains, Revitalizing American Manufacturing, and Fostering Broad-Based Growth: 100-Day Reviews under Executive Order 14017," Washington, DC, 2021.

[14] U.S. Department of Energy, "Interpretation of Foreign Entity of Concern," Federal Register, vol. 88, no. 230, pp. 84082-84102, Dec. 4, 2023.

[15] U.S. Congress, Inflation Reduction Act of 2022, Pub. L. No. 117-169, 136 Stat. 1818, 2022.

[16] J. K. Jackson, "The Committee on Foreign Investment in the United States (CFIUS)," Congressional Research Service, R41916, 2020.

[17] OECD, "Recommendation of the Council on Guidelines for Recipient Country Investment Policies Relating to National Security," OECD/LEGAL/0372, 2009.

[18] L. L. Fuller, The Morality of Law, rev. ed. New Haven, CT, USA: Yale University Press, 1969.

[19] J. Raz, "The rule of law and its virtue," Law Quarterly Review, vol. 93, pp. 195-211, 1977.

[20] L. Breiman, "Random forests," Machine Learning, vol. 45, no. 1, pp. 5-32, 2001.

[21] F. Pedregosa et al., "Scikit-learn: Machine learning in Python," Journal of Machine Learning Research, vol. 12, pp. 2825-2830, 2011.

[22] T. Chen and C. Guestrin, "XGBoost: A scalable tree boosting system," in Proc. 22nd ACM SIGKDD Int. Conf. Knowledge Discovery and Data Mining, 2016, pp. 785-794.

[23] M. T. Ribeiro, S. Singh, and C. Guestrin, "Why should I trust you?: Explaining the predictions of any classifier," in Proc. 22nd ACM SIGKDD Int. Conf. Knowledge Discovery and Data Mining, 2016, pp. 1135-1144.

[24] S. M. Lundberg and S.-I. Lee, "A unified approach to interpreting model predictions," in Advances in Neural Information Processing Systems, vol. 30, 2017, pp. 4765-4774.

[25] R. Binns, "Fairness in machine learning: Lessons from political philosophy," in Proc. Conference on Fairness, Accountability and Transparency, 2018, pp. 149-159.

[26] C. Rudin, "Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead," Nature Machine Intelligence, vol. 1, no. 5, pp. 206-215, 2019.

[27] Chenyu Li, Wenhao Su, and Eric Zhang, “Lightweight Hallucination Firewall for Enterprise LLM Applications: Evidence Consistency, Self-Checking, and Small-Model Detection on TruthfulQA”, JACS, vol. 3, no. 1, pp. 49–65, Jan. 2023, doi: 10.69987/JACS.2023.30104.

[28] Kai Zhang, Siquan Meng, and Eric Zhou, “Evidence-Grounded Trading Desk Risk Memos over SEC Filings: Retrieval-Augmented Generation with XBRL Numeric Verification”, JACS, vol. 3, no. 2, pp. 60–76, Feb. 2023, doi: 10.69987/JACS.2023.30205.

[29] Qiyou Wu, Jingwen Bai, and Xiaohan Zhou, “Evidence-Grounded Financial RAG: Reducing Numerical Hallucination in LLM-Generated Corporate Risk Memos”, JACS, vol. 3, no. 3, pp. 65–84, Mar. 2023, doi: 10.69987/JACS.2023.30306.

[30] Sihan Zhou, Zeyi Li, and Eric Wang, “Evidence-Grounded RAG for Tokenized Trade Receivable Disclosure QA under U.S. Capital Market Standards”, JACS, vol. 3, no. 7, pp. 41–57, Jul. 2023, doi: 10.69987/JACS.2023.30704.

[31] Boning Zhang, Hengning Rao, and Derek Zhao, “Evidence-Grounded RAG for Cloud-Native DevOps: Hallucination-Resistant AIOps Question Answering over Private Operations Documents”, JACS, vol. 4, no. 3, pp. 109–125, Mar. 2024, doi: 10.69987/JACS.2024.40308.

[32] Ge Liu, Chenyu Li, and Eric Zhang, “OpsLLM for Cloud Incident Triage: Bilingual RAG-Based Root Cause Analysis and Alert Summarization for AI Infrastructure Operations”, JACS, vol. 4, no. 4, pp. 97–111, Apr. 2024, doi: 10.69987/JACS.2024.40408.

[33] Chenyu Li, Jingwen Bai, and Samuel Wang, “Evidence-Chain Reliable RAG: Word-Level Hallucination Detection, Source Attribution, and Provenance Explanation for LLM Applications”, JACS, vol. 4, no. 2, pp. 76–92, Feb. 2024, doi: 10.69987/JACS.2024.40207.

[34] Sihan Zhou, Zeyi Li, and Eric Wang, “Long-Document RAG for Contractual and Insurance Clause Analysis in Receivables RWA Structures”, JACS, vol. 4, no. 8, pp. 88–104, Aug. 2024, doi: 10.69987/JACS.2024.40810.

[35] Yunhe Li, “Execution-Feedback and Retrieval-Augmented Generation for Conversational Text-to-SQL: From One-Shot Questions to Clarification-Driven Executable Dialogs”, JACS, vol. 3, no. 2, pp. 1–17, Feb. 2023, doi: 10.69987/JACS.2023.30201.

[36] Siyu Chen, Wenhao Su, and Jacob Ma, “Self-Correcting Text-to-SQL Agent with Error Feedback: A Reproducible Closed-Loop Evaluation on Compact Executable SQLite Benchmarks”, JACS, vol. 4, no. 11, pp. 86–104, Nov. 2024, doi: 10.69987/JACS.2024.41107.

[37] Yunhe Li, “Findable then Explainable: Retrieval–Summary Integration for Code Intelligence on a Lightweight CodeSearchNet Subset”, JACS, vol. 4, no. 7, pp. 65–82, Jul. 2024, doi: 10.69987/JACS.2024.40706.

[38] Shilu He, Haowei Tu, and Isa Liu, “Safe PD Capacity Forecasting with Time-Series Foundation Models and Calibrated Uncertainty for Heterogeneous GPU Clusters”, JACS, vol. 3, no. 4, pp. 48–66, Apr. 2023, doi: 10.69987/JACS.2023.30404.

[39] Daren Zheng, Boning Zhang, and Julie Geibel, “VerifySafe: Toxicity-Safe Agent Responses under Adversarial Prompts with Evidence-Based Self-Verification”, JACS, vol. 4, no. 1, pp. 67–82, Jan. 2024, doi: 10.69987/JACS.2024.40106.

[40] Yuanzheng Chen, Yitian Zhang, David Chau, and Matt Sherman, “Credit Card Default Risk Tiering with Probability Calibration and Uncertainty-Driven Rejection: A Reproducible Study on the UCI Credit Card Clients Dataset”, JACS, vol. 3, no. 4, pp. 31–47, Apr. 2023, doi: 10.69987/JACS.2023.30403.

[41] Daren Zheng and Chenyu Li, “Behavior-Level Jailbreak Resistance via Multi-Stage Refusal + Utility Preservation”, JACS, vol. 4, no. 1, pp. 83–99, Jan. 2024, doi: 10.69987/JACS.2024.40107.

[42] Ziliang Samuel Zhong, Ruiyan Ma, and Hailey Zhao, “Human-Uncertainty Distillation for Calibrated Vision Models on CIFAR-10H”, JACS, vol. 3, no. 2, pp. 77–89, Feb. 2023, doi: 10.69987/JACS.2023.30206.

[43] Daren Zheng, Chenyu Li, and Harvey Davidson, “Continual Red-Teaming for In-the-Wild Jailbreaks via Online Guardrail Updates and Guardrail Distillation”, JACS, vol. 3, no. 2, pp. 35–49, Feb. 2023, doi: 10.69987/JACS.2023.30203.

[44] Jiaying Jin, Tina Huang, and Sam Lu, “A Model-Risk-Friendly Probability of Default Workflow: Calibration, Distribution-Free Uncertainty Quantification, and SHAP Explanations on the UCI Credit Card Default Dataset”, JACS, vol. 4, no. 6, pp. 74–85, Jun. 2024, doi: 10.69987/JACS.2024.40606.

[45] Hailin Zhou and Sarah Zhao, “LLM-Explanation-Enhanced Retail Credit Default Prediction with Gradient Boosting on the UCI Default of Credit Card Clients Dataset”, JACS, vol. 4, no. 5, pp. 102–118, May 2024, doi: 10.69987/JACS.2024.40508.

[46] Jiaying Jin, Tina Huang, and Sam Lu, “Cost-Sensitive Learning, Simulated PU Learning, and One-Class Autoencoding for Extreme-Imbalance Credit Card Fraud Detection”, JACS, vol. 4, no. 6, pp. 64–73, Jun. 2024, doi: 10.69987/JACS.2024.40605.

[47] Yuanzheng Chen, Yitian Zhang, and Matt Sherman, “Going Concern and Bankruptcy Prediction under Extreme Class Imbalance: Cost-Sensitive Learning, Resampling, and Focal Loss with Explainable Financial-Ratio Portraits”, JACS, vol. 4, no. 4, pp. 80–96, Apr. 2024, doi: 10.69987/JACS.2024.40407.

[48] Shenghan Lu and David Zhou, “TinyLLM-Assisted Intrusion Detection for Real-Time IoT Networks”, JACS, vol. 4, no. 8, pp. 72–87, Aug. 2024, doi: 10.69987/JACS.2024.40809.

[49] Q. Xin, Z. Xu, L. Guo, F. Zhao, and B. Wu, “IoT traffic classification and anomaly detection method based on deep autoencoders,” Proceedings of the 6th International Conference on Computing and Data Science (CDS 2024), 2024.

[50] B. Wang, Y. He, Z. Shui, Q. Xin, and H. Lei, “Predictive optimization of DDoS attack mitigation in distributed systems using machine learning,” Proceedings of the 6th International Conference on Computing and Data Science (CDS 2024), 2024, pp. 89–94.

[51] Jiayi Nie and David Zheng, “Ambiguity-Aware HDFS Log Anomaly Detection with Retrieval-Augmented Failure Narratives and Selective Refusal”, JACS, vol. 3, no. 1, pp. 66–80, Jan. 2023, doi: 10.69987/JACS.2023.30105.

[52] Ge Liu, Shilu He, and Isa Liu, “LLM-Augmented Multi-Source Root Cause Attribution for CPU and Network Faults in Microservices”, JACS, vol. 3, no. 6, pp. 39–57, Jun. 2023, doi: 10.69987/JACS.2023.30604.

[53] Jiayi Nie and David Zheng, “Noisy-Neighbor-Aware VM Degradation Risk Modeling with Unsupervised Residual Fusion”, JACS, vol. 4, no. 4, pp. 112–123, Apr. 2024, doi: 10.69987/JACS.2024.40409.

[54] Siyu Chen, Shilu He, and Eddin Sun, “Risk-Bounded GPU Resource Oversubscription via Conformal Demand Envelopes in Production AI Clusters”, JACS, vol. 4, no. 5, pp. 119–134, May 2024, doi: 10.69987/JACS.2024.40509.