Submind YouTube summaries
Thumbnail for Rong Tong: Machine Learning Guided Discovery of Stereoselective Polymerization Catalysts (TSVP Talk)

Rong Tong: Machine Learning Guided Discovery of Stereoselective Polymerization Catalysts (TSVP Talk)

Watch on YouTube

Video summary

Professor Wong Tong from Virginia Tech presented his groundbreaking research on utilizing machine learning to discover stereoselective polymerization catalysts, addressing the critical need for degradable and recyclable alternatives to traditional petroleum-based plastics. His group focuses on polyesters, particularly polylactic acid (PLA), which can be engineered with specific thermal and mechanical properties by controlling the arrangement of chiral centers along the polymer chain through stereoselective polymerization. However, developing these catalysts has historically been a labor-intensive process relying on trial and error, often requiring extensive multi-step synthesis and resulting in small datasets that challenge traditional machine learning approaches designed for organic chemistry with larger reaction databases. To overcome these limitations, the presentation detailed the application of Bayesian optimization, a data-driven method originally used in computer science to tune model parameters effectively. Unlike traditional quantum mechanics-based approaches that rely on specific structural parameters, this framework uses Gaussian process regression to balance the exploration of unknown chemical spaces with the extrapolation of known data. The team developed a specialized workflow involving "featurization" to translate chemical structures into computable descriptors and employed an expected improvement algorithm to propose new experiments iteratively. By integrating synthetic feasibility constraints, such as limiting ligand synthesis to three steps or fewer, the researchers significantly accelerated the discovery process, successfully identifying highly isotactic and syndiotactic catalysts that outperformed random search methods in reaching convergence within just a few iterations. The study further demonstrated how this unbiased approach not only predicts high-performance catalysts but also elucidates the underlying structural mechanisms influencing selectivity. Through feature attribution analysis using SHAP values, the researchers identified key factors such as steric bulk and electronic properties that dictate whether a catalyst produces strong, brittle materials or ductile ones. This mechanistic insight allowed them to design novel aluminum-based catalysts capable of enantioselective polymerization, a process essential for producing expensive stereocomplex PLA with superior melting temperatures and strength. The ability to maintain selectivity even at higher temperatures and molecular weights enabled the scalable production of these advanced materials, which showed performance surpassing standard packaging plastics like PET. Ultimately, this research establishes a robust framework for rational catalyst design that can be adapted to more complex polymer chemistry challenges. By combining global and local descriptors with in-loop analysis, the team created a highly efficient system capable of handling nonlinear relationships and expanding chemical spaces dynamically. The successful scale-up of these processes to industrial-relevant conditions highlights the potential for machine learning to transform material discovery from a slow, empirical endeavor into a predictable, accelerated science. This approach offers a pathway to create new materials with tailored properties while minimizing waste and resource consumption, marking a significant step forward in sustainable polymer engineering.
Read the full video transcript
Um, all right. So, I suspect, oh, sorry. Well, we can start. I suspect that some people from my group will be trickling in as well. Um, but we'll get started today. Uh, it's my great pleasure to introduce Professor Wong Tong who is currently an associate professor at Virginia Tech in the chemical engineering department. Uh, he obtained his undergraduate degree at Budan University. uh went on to do his PhD at University of Illinois Urbana Champagne um and then did some posttocks at both uh MIT and Harvard focusing on polymers for uh biological applications. Um he's been doing some really fantastic great work related to polymer chemistry in general. Um and in particular he's been uh leading the field with using machine learning um in discovering new catalysts for polymerization. So I'm really excited to hear about that. Um and along the way through his work he's won a number of awards um including uh one of the ones is the thim chemistry journals award and also other emerging investigators award. Uh so without further ado I'll hand it over to him. Um but again please join me in welcoming him. >> Good afternoon everyone. Um, thanks Christie for the very nice introduction and it's my great pleasure and my honor to join this GSVP program and also talk about our recent research on using machine learning um for discover the stereo selective um polymerization catalyst. So we are living a materials world and the plastics has been everywhere in our daily life and many of these plastics they are made from uh non-degradable petroleum resources. So they have been used only once or few times and then being trashed and this going to cause huge economic and environmental problems. So people has been proposed to develop degradable and recyclable polymers as alternatives to current non-degradable polyolins. So over the years um people has been develop different types of monomers and trying to make them recycle and degradable polymers and our group has been actively working on this degradable polymer field and we have been focused on polyers. So in this case by introduce the site functional groups or by changing the um mon sequence or changing this polymer topology either linear or cyclic. We have been showing that many of the polymers the polyesters we make uh they have good thermal and mechanical properties comparable to non-degradable polyolifines. So all these um achievement in the chemistry cannot be realized u without the development of the primarization catalyst. So this kind of pummerization catalyst development however it's actually a labor and the time intensive process. It's also um being um based on the try and errors and there's no rational guided um system for how to develop these polymerization catalyst. So just give you an example in this 2021 papers and we actually tried 14 different catalyst and eventually find out uh one of the stereo selective primarilization catalyst working for these reactions. So the question always haunting us is that whether or not we can identify um a process that is unbiased and also a rational guided process to help us discover and also optimize the propriization catalyst discovery. So this kind of search problems is not um only happen in chemical science it also happen in computer science. So basic optimization um that has been um widely used in the computer science is actually a novel approach which can um allow the computer scientists can tune the machine learning model parameters very effectively. So people has been inspired by the success of these kinds of patient optimization approach and trying to use this in the chemistry. So traditionally in the organic chemistry people often use quantum mechanics to compute the uh structure parameters and trying to set up the linear relationship between the structural parameters and their reactivities and however this bas optimization actually take a different approach in this case. Instead of looking at specific quantum mechanic parameters, you look at the data results and trying to balance the exploration of unknown chemical space and also extrapolation of the known data and find out what's going to be the next few minimum experiments and to achieve the global max in this chemical science searching. So in this case the basial optimization has been successfully applied in the organic chemistry field especially in prediction of the reaction yield and also prediction of the reaction in natural selectivities. However all of these reactions has been established based on large uh number set of the re reaction data sets. And the question here is whether or not such basic optimization this data science techniques can be straightforwardly uh applied into this polymer chemistry and this is quite challenging. First is because both organic chemistry and polymer chemistry they actually have different optimization targets. Organic chemistry more focus on yields and this E percentage. However, for polymer chemistry more focus on the molecular weight, molecular weight distribution or the stereo selectivities. Another challenge is that uh for many of the these organic reactions, it has a very large set of the database to work with. For these organic reactions, you can simply change the reactant and to create these high throughput reactions and that can generate large database for people to work with use machine learning algorithm. However, for the polymer chemistry in many cases the monomer and the polymer is actually being fixed. So also in these cases many of the catalyst um that being specifically applied in one of the polymer chemistry also being fixed these required multistep synthesis and often have smalls size database. So this um requires us to develop a highly efficient machine learning algorithm and for the polymer catalyst discovery. So in this case we trying to apply the special optimization in the polymer chemistry and the system we are selecting is actually the polyactic acid. So polyactic acid is actually a very important um polyactors. It's the leading biodegradable materials and it has been largely produced in the industry. It's also degradable and recyclable materials. So one of the interesting properties of the polyactic acid or the POA chemistry is that um this monomer actually have both D and L these chyro centers by arranging these pyro centers along the polymer chain it actually can create PLA with different materials properties the thermal and also the mechanical properties. So the chemistry to doing this um by arranging these chyro centers along this polymer chain is called a stereos selective polymerization. In this case we are trying to apply the b basian optimization for this ring opening polymerization of the lid and the system we are focusing on is this aluminum catalyst systems. So this is because in the literature there has already been 56 different catalyst reported in the literature and many of these aluminum catalyst is actually highly iso selective. Um only a few of them has been shown the ke selective. So the optimization of our goal is trying to find either high PM value which corresponding to this zero block PLA synthesis or the high PR value which corresponding to this hydro hydrotactic P synthesis and our optimization is trying to find this catalyst. In this case, trying to identify the PR value is a selective catalyst is more challenging because in the literature there's not many good data points in this case. So before I show you how we um develop B optimization and the data for using B optimization for the catalyst discovery, I want to walk you through the whole basic optimization process in principles. So on left side actually is hypothetical the search space we're going to work with. So in this case we try to identify what's going to be the maximum data points either for the catalyst with highest PR value or PM value. So the first thing we're going to do is trying to fit the literature data into this chemical space. And uh as as you probably know that the computer cannot like humans eyes directly recognize the chemical structures. So in this case we have to translate um these chemical structures into the language the computer can recognize and uh it is called a featurization process and use this featurization process it can generate descriptors that the computer models can read. And we're gonna try different featurization um techniques and methods and try to fit the literature data in this chemical space. Next, after we fit the data, we're going to do the training. So the model we are trying to establish is called the surrogate model. The survey model is actually the key model that has been used to evaluate each data point in the in this search space and their potential working performance and their potential uh uncertainty variance. So by establishing this search space this can help to train the data and make the prediction. And in this case we're going to use this gausian regression process for the data training. And this sculpture regression process it's actually very effective algorithm and has been shown that it's very useful for the small set of the data. So in this case by adding each new observation into the search space. This Gaussian optimization process can significantly reduce the uncertainty in the whole in the whole search space and help to find out what's going to be the maximum points. So after using this print this survey model um the survey model actually going to generate acquisition function and in this case we use this expected improvement algorithm for the acquisition function generation. So this going to propose a new studies for this acquisition function. New studies in the unknown area what's going to be the next experiment to do to trying to find out what's going to be the maximum optimization data points and uh once it's proposed new experiments we're going to go back into the lab and synthesize these catalyst and evaluate their performance in the polymerization. And once we obtain the experimental data, we're going to feed back into the model to retrain the uh basian optimization model. And this whole process is finished and we call this one iteration or one around. And we hope by doing this repeatedly um we can eventually find out the maximum data points in this whole chemical space. So now I'm going to show you how we develop and discover the serial selective catalyst using this basial optimization framework. So the catalyst as I mentioned we going to work out is this aluminum catalyst with the cellular liant and we're going to divide this catalyst into two parts. One part contain this phenol group and the other part contain this diamine group. And this gonna based on the literature this gonna have 52 um unique substitute lian substitute. And by doing the recombination eventually we're going to have 576 hypothetical uh symmetric lians to work with and to search um which one going to have the zero selectivity. The next thing is that after we have the literary data, we going to find out what's the best visualization method to generate the descriptors for the machine learning models. So we actually use five different featurization strategies for our catalyst and to evaluate their performance, we use this five-fold cross validation. That means for the training data set, we're going to randomly divide this training data set into five parts and the four parts of this training data set going to use the training the data and the one part that's going to be used for the testing. So we're going to evaluate the observed and the model predict uh PM or PR values uh in this case and we find out the DFT moderate and also the EI method they actually have the best absolute errors for the observed PM value and the predicted PM value. So eventually we select this DFT um the density function theory generated descriptors as um the descriptor we're going to work with because the DFT descriptors it's going to provide more structural information eventually for our um mechanism or structure analysis and after you we decide to use the DFT method to generate descriptors we're going to put this literature data into our machine learning model and here we use the combination of the Gaussian regression process and also the expected improvement for the circuit model and acquisition functions and to evaluate their search efficiency. So the result showing that um actually use our basian optimization approach we can quickly reach the convergence which means find the maximum data points within five to seven iterations for either PM value or PR values and we're going to benchmark it uh compared to the random search or the render forest search and all the other um uh algorithms actually have difficulty to reach the convergence even after 10 rounds of search uh iterations. So this proves that our basian optimization has very high search efficiency um in terms of finding out the maximum data points um for the um for the zero catalyst search. So another thing we want to put in the consideration before we propose the new study is the synthetic skills. So because we use the literature data and many of that literature data actually for example for this kind of lian they do multi-step synthesis and in our case we also value the time as uh the time for the synthesis as important because we want to accelerate the discovery process. So in our case if the lians can be prepared with in the three steps and we're going to put the priority to synthesize these types of lian for the um proposed catalyst liant. So here we do three runs of basian optimization search for both PR and the PM values and you can see all these new 33 newly synthesized complex eight of them is highly iso selective and five of them is highly um hro selective so in this case I want to specifically highlight the future selective catalyst synthesis because initially we don't have many good data points and you can see the arrow between the experiment and the predicted model predicted values has huge arrows. However, once we do a few runs of iterations, the errors between the prediction and the experiment has been significantly reduced. Another thing I want to highlight about this process is that um this spatial optimization process is also unbiased. So used the iso selective palace as an example. For example, the A11 based lians in the literature is only show the moderate iso selectivity. However, the model actually proposed two A1 based lians and that's predicted to have high iso selectivities and that's being the true for this iso selective catalyst screening. It's also being true for this uh future selective catalyst and in this case it surprised us is that based on literature this A5 based liant only report once and it has the moderate iso selectivity however the model predicted it's going to have three different tissue selective calis and all of them it's being shown tissue selective so this means the model is actually based on the structure information and make the prediction um for the catalyst search. So this allow us trying to just based on based on our experience to find out catalyst in its model totally based on the structure information and uh trying to avoid missing the potential candidates in the catalyst search and optimization. So after we obtain this catalyst we do the benchmark compared to the literature data and all of these show is the iso selective or hro selective compared with the literature enma and the other thing we want to highlight about is that um these two types of of catalyst actually generate the PLA materials with different mechanical and thermal properties. For example, the iso selective catalyst can generate the serial block PA which has high strength. On the other hand, the hydro selective POA showing in this red curve that has been show high ductilities. So if we combine them together, blend them to make this green curve and in this case it's not only show the strength but also show the ductility and have better performance than the low density polyethylene um for potential applications. So another good thing I want to highlight about this spatial optimization is that not only it can help predict uh a good catalyst it also help us understanding what's going to be the important catalyst features that's affect the iso selectivity or feature selectivities. So model itself that's going to do the feature attribution analysis and to highlight some of the important structure features that we can potentially used in future design and to understand the mechanism. So the algorithm we use for this feature analysis is called the sharp analysis. So they're going to rank all these different descriptors that has been used in the machine learning model and highlight what's going to be the most important features um that's going to affect here is the iso selectivity in this case. So in terms of steric effect you can see the bur volume actually have been very important affect the catalyst iso selectivities and bur volume is actually the volume that surround the mental centers and it's decide the free volume um that another reactant can access these metal centers. So in this case we just look at the bur volume of the catalyst and we find out for the selective catalyst uh PR over 0.98 this has been specifically confined in a small region that has been show selective uh in contrast this is selective callus they actually be in a wide range of the bur volume in this case so this can also look at these electric properties um for different parts of the liant. For example, the homo energy of the diamine part or the homo energy of the uh phenol part. So for example in this case if we fix the diamine part and just lower the homal energy of the the phenol ring part and we can find out it's going to increase the hro selectivity of the catalyst. So this electronic effects also can be very easy to tell by using this sharp analysis. So lastly, we can combine all of these and to generate this linear regression which can highlight what's going to be the most important features and this going to significantly reduce all different parameters we are looking at in this kinds of nonlinear relationships and allow us can potentially future to optimize or design new car. So we have been successfully applied um these spatial optimization optimization strategies for the stereo selective palist um search and we ask ourself can this approach being used in a more difficult problems. So here the problem we want to focus on is another stereo isomer in the POA chemistry and that's the stereo complex. So stereo complex PLA can be made by blend the PLA and the PDLA together. So in the lab it has been show that these types of stereo complex Pa has high melting temperature and also have been shown has the high strength and increased summabilities compared with all other stereo isoblasts. However, because the dactic acid is not exist in the natural environment and the production of this dactic acid and subsequently the dactide has been very expensive, it prevent the industrial production of this stereo complex ta. So theoretically if you want to use the cheap resource to produce this stereo complex ta you have to have this what we call inanto selective polymerization that specifically pick up one of the inantum in this recemic mon mixture and polymerize it inantum unreacted and that's the problem we try to focus on and over the years people has been working on the stereo or indential selective polymerization problems for the lactite and you can see n of them has been very successful eventually when you increase the reaction conversion and it's going to generate the gradient PLA instead of specifically the PLA or PDLA. So fundamentally this is because the catalyst actually mediate a specific mechanism for this indential selective polymerization. So in this case you have to have what we call the inentomorphic side control mechanism. Once the inantum added the monomer added onto the chain end even the stereo error happens this catalyst can still pick up the right enumus and add to the chain end and this kind of micro structure difference you can use the homodoupled proton to tell the structure difference on the other hand um if it's a chain end control magnet that's had been show For most of these POA catalyst, the serial error will not be corrected and the unpreferred monomer has been keep adding onto these chain end eventually that's going to give the stereo block a gradient PLA. And the other thing in the kinetics is that when the preferred monomer is used up the polymerization should stop. However, in many cases um for all these reactions, the polymerization still continue even the kinetically preferred monomer being used up. So fundamentally this is the mechanism challenge how to find out this inomorphic side control for this priorization. So last year we have been reported that we identify this hyro aluminum s aluminina catalyst this biometallic catalyst for the ring opening prization of ala in this recemic lactide mixture and we also identify some of the monometallic archalist can specifically select the DA in this recemic lactide mixture. So the question we want to address is can we use the basian optimization and uh find out some of the model metallic as aluminum catalyst and uh for the inential selective primarization synthesis. So directly transfer the previous spatial optimization model to this new to this new problem is challenging. So the first is that in our previous model when we make the prediction we only focus on PM or PR values. However, in this case, not only we need to focus on the primarization in selectivity, but we also need to pay attention to the conversion and trying to avoid the gradient polymer PLA productions and the other challenging part is that in the previous system there's already literature data points that's 56 unique different catalyst and however for our case there's no pre-existing data sets. uh we can refer to. So this for the challenge uh to us is that we have to establish a very highly efficient basian optimization framework for this indential selective catalyst discovery and after the optimization and here is our modified or improved basing optimization framework. So we still use the DFT to generate the descriptors. However, just different from the previous where we do the fragmentation use the DFT calculate each liant parts. We also calculate the whole palace part and also the specific substitute group in each of these liant um substitute. So we create what we call this global local combined DFTbased descriptors. For each catalyst we have generate over 158 descriptors. And this going to provide rich chemical information for the facial optimization the machine learning model to study and to train and to uh predict what's going to be the highly to selective catalyst. And the other improvement we made is also the in loop analysis. So different from previous approach we do the sharp analysis after all the patient optimization search finish. In this case, after we predict the catalyst doing the experiment in the lab and then we immediately do the results analysis and try to see what's going to be the important features that affect either on the alpha values or on the conversion values. So this allow us to add new liant component into this chemical space to targetly expand the chemical space and to improve the search efficiencies. And next we're going to show you how we apply this improved framework for the inventor selective catalyst discovery. So we focus on this Sbased aluminum catalyst and in all these cases these kinds of catalyst can be directly um prepared with these three steps and our target is that we want to find out the RF the polymerization in n selective values over 0.9 and the conversion that's going to stop around 50%. So we first um compare our featurization method to the previous method or the whole paral featurization method and we show that use our combined global local descriptors the arrows between the prediction and the obser observe has been significantly uh decreased and also it provide more informations for the magnetism studies and this has been shown that in this observed and predict polity plot and this training and test they have low arrows and next sorry next thing I'm going to show that this improved the basian optimization search they can reach the convergence to find out alpha values within three to four runs and this is um much efficient than the random search or another SMAC algorithm that has been shown in the literature also effective for the machine learning and our improved the basian optimization has been showing highly efficient to reach the search convergence. So when we perform the round one prediction and testing and we look at um our results and this figure shows that the primarization results and you can see in our initial data sets and the round one data sets some of the primarization didn't have any of the inential selectivity some of them have low reactivities and some of them even produce undesigned gradient PLA so the primarization results has been widely distributed and also in this chemical space showing that most of the prediction um or the round one search has been confined in a small region in this whole chemical space. So we want to expand targetly expand this chemical space and also improve the search efficiency of the alpha and conversion. So we perform the sharp analysis and in this case after the first round in this case we're going to show that the R1 group in this A part it has to be less stereial bulky um to have very good conversions around 50% for thisization and also you need to change the structures on this bapial part on this um B part so We then put targetly put um the A part with less bulky iron group and also add new uh diamine groups into the chemical space and to see if whether or not these kinds of lian can really improve um the search efficiency. And indeed we found in this round two and round three um the difference between the experiment value and the prediction has been significantly reduced. You can see from this plot and also all the experiment values in the run two and the run three predictions they have been close to 0.9. So over these 51 newly synthesized catalyst 28 of them has alpha over 0.9 and the conversion is reach about 50% and that's the goal we want to achieve. You can also see all these whole chemical space searching and round two and round three has been widely explored and to try to find out what's going to be the most uh efficient uh in selective catalyst candidates. So we also performed the NMAR studies trying to confirm these catalyst do follow this inomorphic site control mechanisms. So in this case we duterate label the massive groups in the LLA. So by doing so we can tell the reactivity between the LLA and the DLA. So if there's a peak showing up or changing in the massive region in the proton and the M that's indicate the DA mass group has been involved in the primarization and we monitor the whole pmerization process. We didn't see the DLA has been polarized in this case. So that's confirmed that our lead palace did follow this inomorphic site control enchantment. As a negative control, we use this chain and control mechanism aluminina catalyst and you're going to see in this case the DA going to be involved in the primarization and you do see the peak that's showing up in this region. And we also monitor the kinetics of the whole primarization process. Initially they have been showing the first order kinetics. And once the conversion reach about 48 to 49% you're going to see the polymerization almost stop. This also confirmed that our paralyt actually had this uh inentto selectivities and going through inomorphic site control inment. So finally we do the sharp analysis for both alpha and conversions and uh we also do the clusting of different kinds of uh descriptors because in this case we have over 100 different descriptors and we want to avoid the redundancies in the uh analysis of which types of descriptors affect the alpha or conversion values and we use that to establish the linear models in this very complex nonlinear relationships and to highlight some of the most important features that's going to affect either the alpha values or predicted conversion values. So lastly I want to highlight um some of the unique features of our lead catalyst and during the basian optimization process we found out that some of our lead catalyst actually didn't have significant reduction of the inential selectivity even we increase the temperature over 100 degree. So this is very interesting and unique in terms of inential selectivity. Usually when you increase the reaction temperature these types of inential selectivity or stereo selectivity going to be decreased once the temperature is high. However in our case this field lead catalyst didn't have these kinds of inential selectivity reduction. So this allow us to produce this in selective pmerization in the industrial relevant B primarization conditions. So here you show that we did this uh bulk polymerization uh in a 10 grand scales. We we just mix the monomer with the lead catalyst and heated up over 150 degrees. So the lactide going to be melt and polymerization starts. So after the polymerization stop we can separate the PLA and also recycled the unreacted monomer which going to be have high E values enriched in the dactic acid. So the homodoupled NMR and the C30 NMR confirms that these types of polymerization did have the indential selectivities. Another interesting point is that when we increase the feed ratios of lactic acid to the aluminum catalyst, we're going to see we can control the molecular weight increase linearly. And this um blue points showing up shows that the native selectivities did not decrease even we have molecular weight over 100k in this book polymerizations um condition. um with that strategies we can use this uh in selective polymerization actually to produce the stereo complex POA. So in both cases we can use either S or R colorless to pro produce this highly isotactic P and the PDA and mix them with the batch one stereo complex PLA. So the inenttoriched unreacted model can going to go through this bulk ring opening priorization and produce another um highly isotactic PDLA and the PLA and mix them to produce another set of stereo complex PLA and this going to highly efficiently use up the monomers for the stereo complex um PLA productions. In this case, we see that both two batches show the increased melting temperatures. In this case in the uh DSC analysis in the mechanical property analysis we found out actually the batch one stereo complex PA not only showed the high strength but also show the improved ductilities compared with all other stereo isomers and this is even better than the current um standard packaging materials the PET and this has been showing this uh maybe the introduce of a little Steer arrow along this zero complex polymer chain. Uh not decrease the ductility but increase or not decrease the strength but increase the ductility by changing the flexibility of the polymer chain and we still um doing the studies trying to see what's going to be the reason for this improved adaptivities. So given this easy scale up method and also the high if excellent performance of this uh serial complex PA we believe such method u can have industrial relevant um productions for future applications. So overall the message I want to um send to you is that so by developing a highly efficient basian optimization models and this going to be uh have huge potential for the polymer chemistry. Not only it provides an unbiased and rational guided way for the polymer catalyst development but it also help to finding out interesting structure features and help us identify new catalyst and also for future catalyst design and uh I want to thank uh my group my students doing this work and the initial um the basial optimization framework is through the collaboration with another group in Virginia Tech um and also Shiao and that's the reason follow up to develop this basian optimization for the in selectiveization uh catalyst and the founders from NSF and ACSP and uh thanks for your attention and I would like to answer all these questions. is >> great. Uh thank you so much for the excellent presentation. Uh any questions for Rang? >> Uh thank you for your your presentation. I wanted to ask them, you calculated a huge library of descriptors for different catalysts and this saline type type catalyst can be active in polymerization of many various monomers not only PLA right >> but it cannot be used for bio optimization in this step because it doesn't have data set of um this catalyst applied in polymerization of this Yeah, I I think this is a very good question. So um it's actually a problem we also want to work with is that whether or not uh what type of like these types of knowledge can be translatable to other polymerization system even it's very close for example um other kinds of like the beta lacone and also use this um cell aluminum catalyst or similar types of catalyst whe whether this can be translatable and the current is we don't know. So the thing is you have to establish um some kinds of um initial data set like you need to directly apply the catalyst to test on the new model systems and then know what's the steer selectivity data and then to further find out whether that's going to be help you to find out a new callus. So I would say this is case by case and this is also being true for many of the organic reaction uh machine learning optimization problems. One of the like the coupling reaction you predict the yield follows a specific model may not be directly translatable to another kind of coupling reactions that's even use the same metals for example nickel metals. So all these actually you have to have um current understandings you have to have some initial data set to start to work with and then see if the knowledge can be translated. But can this knowledge be used for creating like library of various catalysts mostly various I mean um the degree of v variety between these catalyst might be the maximized right and can the library of catalyst be created that some um experimental groups can test this and create data set and then you can use >> yes I I think that definitely it is I um it's not um related with this project. I think um in organic chemistry they actually have a force liant and people create these types of libraries. Um also the amin groups and I think the sigma group is also creating these libraries for the organic reactions doing this coupling reactions. So definitely you can do the DFT computation of different types of P list and generate a library that contain all these descriptors and see maybe one day it can be used for um a new reactions or for a new specific applications. >> Thank you. >> Very good. >> Any other any other questions? >> Any other questions? Yeah. >> Yes. Um, very cool stuff. This is very excited. Exciting. I'm also looking into something. So, it's fun to to see. Um, I've written down slide 14 and 31, but I think it's more of a general question. Uh, 14 or 31 because you have a whole set of different lians that you look into or that you use as a as a data set. 14. >> Yeah, it's just a more Yeah, >> this is >> Yeah, >> this is basically the iteration process. >> Um, from a more general standpoint, how many of the suggested catalyst synthesis were chemically not possible? So of course now you only show or the ones that are shown are the ones that you know had an outcome let's say. >> Yeah that's right. So um in our initial design um use the literature data like some of the ligans actually take five six steps to make >> and we think that's impossible for us. >> Yes. to repeat that uh multi-step synthesis and that's why we introduce the synthetic scale and try to find the synthetic steps within three steps. So and also during our search um some of the aluminum palace liant and that's just we going to make the liant and that's the easy part and then when you add the trimester aluminum into the liant and uh in some rare cases um maybe the steeric group is too bulky >> and these kinds of catalyst it's very difficult um to work with. They have very low solubilities or you have to heat up the reactions um at 100° to make it soluble. >> So we do have the difficulties. >> Yes. Um in terms of the synthesis [clears throat] >> but >> you can see our previous paper we do report it like machine learning algorithm they predict this and we make this and this has low solubility and that's [clears throat] why possibly the yield or the serious selectivity value it's low >> as you it's the rare cases >> it's do going to happen once you make this hypothetical um careless leg predictions happen. >> Majority is >> yeah majority of them is quite robust and straightforward to prepare. >> Cool. >> Your five-fold validation Yeah. >> was the same across all of the feature sets. >> Yeah. >> And do you think that was evenly distributed across chemical space as much as you can put chemical space on like an ordinal scale? or >> so this is not um for the whole chemical space. This is for just the training data set, right? You just randomly divide it into >> five parts and we not run this just once. We run multiple times and so that's why you see the arrow the standard arrow bars in this um plot. >> Yeah. So >> uh so do you feel like moderate or EI? >> Yeah, moderate EI they they different yeah different visualization in this case. Um we think the initial training data set is not a lot. It's only 56 data points. So you're going to see the arrow probably similar to the DFT. But the advantage for the DFT is that when you do the sharp analysis, it can tell you which feature is important and later on you can do the magnet studies or like later on like the second patient optimization models we can um see what's going to be the important features and we can add leg into the chemical space. >> Okay, so that's the >> that's the reason we use it's not a black box. I have a more kind of general question. So I mean you still had to make make 50 lians each with a three-step synthesis. So your poor students run 150 reactions in order to make the lians and did the polymerizations like are you looking into automated synthesis as well or not at this stage? Yeah, not at this time. Um that's very good point the automated synthesis and it's actually easy to do in the organic reactions like um doyas work and also how's work and they actually apply this high finger into the crossoupling reactions and in our case we have to prepare the liant and uh this is like three step is the maximum many of them it's actually can be done with in one to two steps. Um and once you create this liant, the difficult part as we just discussed is add the aluminum into together with the ligan and make the catalyst. So that's the most difficult part and each time we cannot guarantee you do these kind of mixture and evaporate the solvent you're going to get the right calis structures. it may have the mixture and they have been reported for this aluminina they could have this dma aluminum with the salon liance and that's what we trying to avoid it um in this um catalyst synthesis and we want to make sure it's monometallic in the structure so that's sort of like the limitation step for us to prepare the high stud >> great um any other questions or online. Yeah. Okay, we're okay. Okay. Well, um if that's the case, we'll wrap up now. Thank you again for a really insightful uh presentation and I hope that many of you in the room will use this opportunity to come and speak to uh R um about machine learning and so on. Great. So, thank you