Goto

Collaborating Authors

 pursuing


ABKD: Pursuing a Proper Allocation of the Probability Mass in Knowledge Distillation via $α$-$β$-Divergence

arXiv.org Artificial Intelligence

Knowledge Distillation (KD) transfers knowledge from a large teacher model to a smaller student model by minimizing the divergence between their output distributions, typically using forward Kullback-Leibler divergence (FKLD) or reverse KLD (RKLD). It has become an effective training paradigm due to the broader supervision information provided by the teacher distribution compared to one-hot labels. We identify that the core challenge in KD lies in balancing two mode-concentration effects: the \textbf{\textit{Hardness-Concentration}} effect, which refers to focusing on modes with large errors, and the \textbf{\textit{Confidence-Concentration}} effect, which refers to focusing on modes with high student confidence. Through an analysis of how probabilities are reassigned during gradient updates, we observe that these two effects are entangled in FKLD and RKLD, but in extreme forms. Specifically, both are too weak in FKLD, causing the student to fail to concentrate on the target class. In contrast, both are too strong in RKLD, causing the student to overly emphasize the target class while ignoring the broader distributional information from the teacher. To address this imbalance, we propose ABKD, a generic framework with $α$-$β$-divergence. Our theoretical results show that ABKD offers a smooth interpolation between FKLD and RKLD, achieving an effective trade-off between these effects. Extensive experiments on 17 language/vision datasets with 12 teacher-student settings confirm its efficacy. The code is available at https://github.com/ghwang-s/abkd.


Much Younger Men Keep Pursuing Me Online--and They Have the Same, Startling Fantasy

Slate

Feeld Notes is a column about a middle-aged woman who suddenly realizes she wants to have sex again--and the beguiling app she uses to do it. Not to imply I get liked by a lot of men--ha!--but of the men who do like my profile, a significant percentage of them are substantially younger than I am. Listen, I'm not a cougar, an appellation that, by definition, suggests a certain predatory instinct. As I've explained previously, I don't have it in me. But I'd be lying if I said that that some part of me isn't delighted to think that the pictures and the words on my profile project a sort of youthful exuberance.


Pursuing a Passion for Machine Learning

#artificialintelligence

This story is part of an ongoing series in which we highlight graduates of Capital One's Machine Learning Engineering Training Program (MLETP), a 160-hour program that teaches software and data engineers the skills necessary to work in machine learning and AI. Pradeep picked up the value of continuous learning from his mother, who earned multiple master's degrees and a Ph.D. in education. So after becoming a software engineer at Capital One in 2017, he was quick to embed himself in our culture of growth and development. Pradeep followed his curiosity and began developing skills in machine learning, a form of artificial intelligence that can automatically predict outcomes. Capital One uses machine learning to create real-time and intelligent customer experiences that bring simplicity to banking.


What Artificial Intelligence Startups are Pursuing

#artificialintelligence

Artificial Intelligence is a niche and fragmented Market and has greater potential in driving future Business Decisions. Artificial intelligence is clearly on the mind of the biggest companies in tech. Of note, Facebook CEO Mark Zuckerberg recently called the development of artificial intelligence one of his top three key goals over the next 10 years. IBM launched a $100M Watson venture fund in January. And Google has already acquired a handful of AI startups including DeepMind Technologies of the UK for $500M.