Sunday, September 6, 2020

ניצוצות -- קידושין -- כסף, שפת הכסף, סקאלה של הכסף

הקבוצות נקבעות על ידי המרחק היחסי בין האלמנטים, הפרטים, במדד מרחק מסוים.

הסקאלה נקבע על ידי מעבר פזה של הסדר/אנרגיה במערכת

כאשר יש מדד אחד, דאטא שהוא בממד אחד, אפשר לתאר את הריחוק בקלות, מה קורה כאשר יש שתי מדדים.  האם נכון לקבוע כל הקבצה במדד שלו ואח״כ לראות את הסקאלה של המערכת או לקבוע את ההקבצות על בסיס הדו-מימד ורק אח״כ לראות את הסקאלה?

קידושי איש, יש מדד של שעיבוד, החתן מתחייב על פה לשאר כסות ועונה.  במערכת ההתחייבויות יש עוד דוגמאות של התחייבות בעל פה, וזה המשך הסוגיא, התחייבות של עולה (קורבן) התחייבות של מלווה וגדר שיעבוד.

יש עוד מימד בקידושין, מימד הקהל, הידיעה המרחבית שהאישה הזאת אסורה על כל אדם אחר פרט לחתן.

אם המדד היא פרוטה, דבר השווה לכל נפש, השאלה היא מי בתוך המערכת של פרוטה, הקהל של העיר (רב יוסף כל דהו), המדינה, האומה (יחסית לאיסר האיטלקי).

אם המדד היא דינר, דבר שאינו שווה לכל נפש, היא מבטא שפה ׳חצי פרטית׳ בין החתן לכלה, היא מבטא את התחייבות הערכית, העצמאית של אישה, אישה מכבדת את עצמה

כאשר הגמרא אומרת מלווה על פה מדרבן אינו גובה מהלקוחות, משום שאין לזה קול, היא משלוות בין שתי המימדים האלו.

Sunday, August 23, 2020

עירובין דף י״ד -- הרב דב ברקוביץ

אני מקשיב ל-  https://www.youtube.com/watch?v=CoFI8fEcZKQ

הרב בקרוביץ שואל כיצד פועל הקניה על ידי אמירה לגבוהה?

ונראה לי שכאשר קיים אצלי שתי זהויות שונות, נניח הזהות הפרטית והזהות הציבורית, זהות אחת יכול להקנות לזהות שניה.  וזה רק יהיה אפשרי אם קיים ניוק מלא בין שתי הזהויות, דבר שאינו ממש בריא לנפש האדם (סקיזופרנה)

-=-=-

עוד למדתי בדף היומי, מסכת עירובין דף י״ד, הפתיחה מלמדת על מבנה, הקורה צריכה להיות חזקה ולהידמות להיות חזקה כדי להחזיק אריח, כלומר שאפשר לבנות עליה (רש״י), ואולי הכוונה היא שמסתכלים על המבנה מבחוץ, ככלי לדבר אחר, ולא התוכן.   מבוי היא מקום הילוך (למדנו בתחילת המסכת, ולא כמו סוכה שהיא לדיור), אבל הקורה הופכת אותה לבית שיש לא תקרה, מקום ישוב.

ובהמשך, הים של שלמה, שיש לה שפה עליונה עגולה ובסיס מרובה, ולי זה מזכיר מבנה של אדם, ראש עגול וגוף מרובע (קצת קבלי זכור לי), ואולי גם מזכיר מבנה של מבוי, עליון שונה מתחתון.

=-=-

ומזה אני למד שיש שתי זהויות העליון והתחתון, והגדרה של העליון משפיע על הגדרה של התחתון.  וצריכים להגדיר בצורה חיצונית, על ידי יכולת לקבל אריח, על ידי צורה שונה, עיגול ולא מרובע, להפריד בין שתי הזהויות האלו, ואז אפשר *מבלי שהשפע ירד מאחת לשניה* להגדיר את הזהות של השני יחסית לראשונה.

Tuesday, July 21, 2020

on models - curse of dimensionality - scale dependence

Pandemic models are being focused on now.


Typically models have parameters.  Then the discussion is how to learn the parameters from the data.

In general there seems to be a scale of models, from those constructed solely on a priori knowledge (expert systems) to those constructed solely on data.

For example, a model of an epidemic may have parameters like, RO, .... Then given real data, the parameters are estimated and the models is tested.

Alternatively, we can start with data, and try and create models that represent that data.  However, keep in mind that the choice of data in itself is a bias introduced by a priori knowledge.

So in truth there is no difference between these two approaches, just a matter of degree.

Given lots of data there is a tendency towards the later side of the scale, given little data the shift is to the former side of the scale.

The curse of dimensionality is a good metric to help decide which side of the scale is reasonable, to define 'lots' or 'little' amounts of data.

However, often overlooked in the nature of the curse of dimensionality is scale dependence. See my blog post on the topic.


So what are we measuring and what are we predicting.  Are they on the same scale?

So there are two types of problematic statements.  Statements that are invalid: a triangle with four sides or a wheel of a car that has a destination (cars have destinations, wheels have rolling characteristics)  Statements that are not supported, by mixing levels they hide the curse of dimensionality and hence the number of observations needed to validate the statement.

mixed scales:
A pandemic model that has the virus 'intent', a virus utilizes airplanes to help itself reproduce.

Mortality is a function of hospital beds (higher level scale) and viral load (lower level scale).

=-=--=
from : https://medium.com/amnon-shashua/can-we-contain-covid-19-without-locking-down-the-economy-2a134a71873f

Let m be the size of the low-risk group and let ν be the probability that a person that comes from the low-risk group will develop severe symptoms, assuming the person is currently sick
...
p* be the current, unknown, percentage of positive cases among the low-risk population and let k be the number of severe cases among the low-risk population from today until one week from now.
...
To do this, we will sample n persons, uniformly at random from the low-risk population, and will derive a lower bound on p* based on the number of people that came out positive. "
=-=
Seems to me all the variables and analysis are at one level - that of an individual patient, either a person is in a low-risk or high-risk group based on age.  A person tested positive or not.  A person is severely ill or not.  The observations are of individuals.  Thus the entire statements computes.

compare to:

https://www.statnews.com/2020/03/16/coronavirus-model-shows-hospitals-what-to-expect/

CHIME (“Covid-19 Hospital Impact Model for Epidemics”), built by Penn’s Chivers and others in “predictive healthcare,” is a basic epidemiological tool of infectious disease spread called a SIR model. It takes what’s known about the number of susceptible (S) people in an area (which for Covid-19 is everyone, since no one has immunity to the new coronavirus that causes it), the number of infected (I) people, and the number of recovered (R) people (who are presumed to be immune from subsequent infection). Because of the disastrous rollout of Covid-19 testing in the U.S., the researchers assume that only 15% of cases have been detected (but say it could be even lower).
The model then uses the best current estimates of how long someone is infectious (14 days); how many new cases each infected person causes (called the effective reproduction number, it’s about 2.5); the percentage of Covid-19 patients who need to be hospitalized (5%, reflecting the fact that most people have only mild or moderate illness); the percentage who need to be in an ICU (2%) or on a ventilator (1%); and the length of stay for each of these three.
-=-=-
and to:
https://www.statnews.com/2020/02/14/disease-modelers-see-future-of-covid-19/

The computers that run disease models grind through calculations that reflect researchers’ best estimates of factors that two Scottish researchers identified a century ago as shaping the course of an outbreak: how many people are susceptible, how many are infectious, and how many are recovered (or dead) and presumably immune.
=-=-
the Scottish researchers limited to three variables all about individuals virus reaction.
However, note how the CHIME model includes the 'reproduction number', which is at a different scale, that of community relationships.  And it also includes ICU usage of a patient which is a function of viral load, how much virus was communicated.

=-=-=
Now it is very possible that there are other factors to the Hospital Load that are not taken into account by the model and its parameters:

https://www.straightdope.com/columns/read/1734/is-there-an-anti-placebo-effect/
While Krenztman uses the term “placebo” for both positive and negative effects, “nocebo” is finding more use these days. As you might have guessed, the nocebo effect is the opposite of the placebo effect. In Latin, nocebo, which only showed up in English usage in the last decade (and, in fact, is not even recognized as a real word by my word processor’s dictionary), means “I shall cause harm or be harmful.” While the medical profession recognized a while ago that they needed to take into account the placebo effect, they have only recently recognized that they need to also take into account the nocebo effect.
Like the placebo effect, the nocebo effect is usually generated by “beliefs, attitudes and cultural factors” (http://quinion.com/words/turnsofphr ase/tp-noc1.htm). This occurs when the expectation of deterioration is created. For an extreme example, the July 1997 Harvard Mental Health Letter notes that the nocebo effect has been credited with causing “so-called voodoo deaths.” In other words, people who truly believe in voodoo and believe they have been cursed by a voodoo practitioner may be so affected by the nocebo effect that they actually get sick and die. The article further notes: “For surgical patients, the expectation of death on the operating table can be fatal. In one study of people with asthma, deliberate misinformation about the effects of medication reduced its effectiveness by nearly 50%. Also, allergic reactions can be induced merely by telling the patient that they are receiving a substance to which they are allergic, when in fact they are receiving salt water.”








Truth -- Dr. Shermer's perspective

listening to:
https://www.skeptic.com/skepticism-101/what-is-truth-anyway-lecture/?mc_cid=4faf75f4bc&mc_eid=a51253749f

What I hear so far is that truth is determined by proving a causative relationship between the elements that constructed the truth.  Something is true because we understand how it got there.

I don't understand why 'truth' should be related to understanding causative relationship.  Or said differently, as M. Shermer himself says, since correlation is not causation and there are many confounding variables, it is very difficult to prove causation.  I would argue, impossible to prove causation, just different degrees of confidence/Information/energy in a system that provide for causation.  Hence, there is never 'absolute truth', and without that what is the point of a definition of 'truth'?

Rather, I would say that a feature of truth is that it must be communicated.  Perhaps a feeling need not be communicated to others, but a truthful fact is only interesting if it is objective truth. Hence there are at least two people, to remove the subjective element of the fact.

If we agree that any definition of truth includes the ability to communicate that truth, then we should agree that truth like language requires an agreement between parties.  Hence truth has nothing to do with causation and everything to do with mutual consent.

Causative relationships are good tools for convincing arguments, and help build mutual consent.  But in of itself a causative relationship is not 'true'

Monday, June 8, 2020

מספרים, כמות ואיכות, בהלכה והלכה למעשה

Is this a good example
https://ashlag-cause-and-kook-affect.blogspot.com/2018/03/scale-dependent-halacha.html

----

Two conversations:

I)  What justifies the use of force by the government in dealing with COVID?
--
1. The government needs to protect individuals (citizens), anyone can infect a person, hence the government can limit the freedom of the individual for the good of others 
2. COVID is a pandemic (epidemic) it affects the entire country, hence it is a national crisis and permits the national government to step in.

Or the conversation is:

II)  Why does the Mishna say the renter is absolved from paying his debt in a national crisis? (Halachic question today, do you need to pay rent for a store that the government shutdown)
--
1. The renter has no remedy, אונס, but if the renter could solve the problem (like by bringing water from the main river for himself) then he is required to pay his debt
2. The renter had a contract with certain expectations, natural events, COVID is unnatural hence it is not covered in the contract (similar to (1) in that a form of אונס)
3. The renter can say to the owner, you would have lost rent either way, since no one would pay you
4. The renter is only obligated to pay when he is the owner (private property, even if temporary) of the property, ownership is a manifestation of a private system, the national crisis removed the ownership, by removing the private system, both the renter and the owner of the property are now in the same system, the national system, and neither own any property 

III) What happens if a large river dries up, yet the farmer can solve the problem?  Or is the definition of a Makat Medina that which can't be solved?
--


בבא מציעא ק״ד
מתני'
 
המקבל שדה מחבירו והיא בית השלחין או בית האילן 
יבש המעין ונקצץ האילן אינו מנכה לו מן חכורו 
אם אמר לו חכור לי שדה בית השלחין זו או שדה בית האילן זה 
יבש המעין ונקצץ האילן מנכה לו מן חכורו:

גמ' 
היכי דמי 
אילימא דיבש נהרא רבה אמאי אינו מנכה לו מן חכורו 
נימא ליה מכת מדינה היא 
אמר רב פפא דיבש נהרא זוטא 
דאמר ליה איבעי לך לאתויי בדוולא 

[מה כוונת רב פפא, האם הכוונה לומר שיבש נהרא זוטא בוודאי אינו מכת מדינה, אבל בכל זאת נשאר מצב שאינו לפי הציפיות, לזה מתרץ רב פפא, הציפיות הוא שאיכר יעבוד קשה.  או הכוונה היא שההגדרה של מכת מדינה היא סוג של אונס, שאינו יכול לתקן בעבודתו]

אמר רב פפא הני תרתי מתניתא קמייתא משכחת לה בין בחכרנותא בין בקבלנותא 
מכאן ואילך דאיתא בקבלנותא ליתא בחכרנותא ודאיתא בחכרנותא ליתא בקבלנותא:

תוס׳ דאפשר לאתויי בדוולא. 
וא"ת אפילו לא אפשר נמי אמאי מנכין לו דכיון דלא יבש נהרא רבה תו לא הוי מכת מדינה 
[מה ההגדרה של מכת מדינה?]

וי"ל דנקט האי טעמא לאשמועינן דאפילו יבשו נמי שאר יאורי שדות אחרים דהשתא הוי מכת מדינה אפילו הכי אינו מנכה לו כיון דאפשר לאתויי בדוולא מנהרא רבה דדוקא באכלה חגב או נשדפה שאין יכול לתקן הקלקול ע"י שום טורח מנכה לו 
[בתחילת דבריו היה נראה לי לומר ׳אפילו יבשו שאר יאורי שדות -- לא הוי מכת מדינה, זה שהרבה יחידים סובלים אינו מגדיר מכת מדינה -- אבל לא כך מפרש תוספות]

אי נמי יש לומר דודאי בחכירות לא היה צריך לטעם זה אפילו בלא אפשר לאתויי אתי שפיר אבל משום קבלנות איצטריך שלא יוכל לומר לא אתעסק בה ואפ"ה לא אשלם במיטבא לפי שאינו מחוייב להביא מנהרות אחרים הרחוקים ביותר ועל כן הוצרך לומר דלא יבש נהרא רבה ומצי לאתויי בדוולא ולכך מחוייב לעשות בקבלנותו:

=-=-=-=-=-=--=-


- הרב רפפורט
ובכן, הנה הדבר: יש כמה וכמה מעברי פאזה הלכתיים לגבי אנשים, הקשורים במספרים, כמו בין תשעה אנשים לעשרה שהזכרת). כמו כן יש בית של שלושה, של עשרים ושלשה ושל שבעים ואחד, ולכל אחד סמכות שונה לגמרי. יש קהילות, ובקהילה מקבלים הכרעה על פי רוב. כל מעבר כזה צריך דיון בפני עצמו. בכלל לא ברור אם יש כלל אחיד מה מודדים בקהילה כדי לקבל החלטה, ויש בזה מחלוקות.

עם זאת, יש מעבר פאזה ברור לגמרי כאשר מדברים על הציבור כולו (כלל ישראל - לא רוב), יש לזה השלכות רבות, כמובן, אבל אזכיר כאן השלכה אחת. אינני מדבר על הציבור מבחינת אחריות שלטונית, אלא מבחינת קבוצה של בני אדם (ששים רבוא בתחילת הווצרותו של עם ישראל). המתרחש בכלל ישראל, דומה, באופן פרקטאלי גם למתרחש בקהילה ובמנין אנשים, אבל כיוון שהסקאלה שונה, גם המהות שונה.
בציבור של כלל ישראל הכהנים עובדים בבית המקדש עבור כולם. לא במובן של חלוקת תפקידים, אלא כל ישראל נמצא במקדש למרות שרק מעטים נמצאים בפועל. היחיד בתוך ציבור מגלם (personifies) את הציבור כולו. לכן כל אחד נמצא במקדש (ע"ע מעמדות), כל אחד עובד את אדמתו, ובתחילת הכיבוש כל אחד גילם את כלל ישראל כאשר כל שבט הלך וכבש לעצמו נחלה לאחר מות יהושוע (עי' רמב"ם תרומות א, ב). לכן מצוות השייכות למצורעים, לזבים וכו' שייכות לכלל ישראל. "כולנו מצורעים"... והמצוות האלה נחשבות שוות בכל. "שווה בכל" פירושו נמצא בפאזה הציבורית כאשר אחד מגלם את כולם. הסיבה לכך היא תלמוד תורה. כיוון שכל גבר חייב ללמוד את כל המצוות, הוא שייך לכולן, כמו שראינו בדברי "אגלי טל" בפתיחה. כל דבר שאינו שייך לכל הציבור (אפילו אם אינו שייך לנשים) לא מתקיים בו הרעיון שאחד מגלם את כולם. לכן כאשר גבר מגלח את פאת ראשו, לא "כולנו" מגלחים את פאת ראשינו, כיוון שנשים אינן בכלל האיסור, וגם אינן חייבות בתלמוד תורה. אבל כל מצווה השייכת בנשים, אבל רק למספר קטן של אנשים ונשים (כמו אלמנה לכהן גדול) שייכת לכולם "כולנו כהנים גדולים שאסור להם לשאת אלמנה". נדבר על זה בעזה"י ביום שני. כך נוכל להבין היטב את ההבחנה מתי עשה דוחה לא תעשה ומתי לא.


=-=-
בוקר טוב, אחרי השיח אתמול וקראתי שוב את המחשבות של הרב פה, נראה לי שהרב אומר שקיימת שתי סקאלות עקרוניות, ציבור ופרטי, אבל קיימות גם מרחבים ביניהם, אלא שהסקאלה של המרחב היא נמדדת על בסיס שתי העקרונות, כאילו האנרגיה של הציבור הכללי נמשך לתוך מרחב קהילה (ציבור קטן).  ואני טוען שקיימות סקאלות ומרחבים שונים, וכל אחת יוצר את האנרגיה שלה מתוך התכונות של הפרטים

=-=-=
- הרב רפפורט
אתה מתאר את מה שאמרתי נכון. על הרעיון שלך חשבתי גם כן, ואביא אותו בפעם הבאה כאשר נדון בקהילות בתוך עיר אחת. זה יהיה למחרת החתונה שלכם... שבת שלו' כל טוב ומזל טוב

=-=-
אני חושב קצת על הדיון הבוקר, נראה לי שהרב הציג נושאים/ערכים שונים בסקאלות שונות, לדוגמא אם נמדוד גודל קהילה, אולי קהילה גדולה יכולה לפסוק לגבי חֵלב אבל קהילה קטנה רק יכולה לדון לגבי צבע הסידור.  ואני חושב שככל שגדל הקהילה הערכים הם יותר כללים, ולכן יותר מופשטים,  קהילה גדולה מאוד יכולה לפסוק לגבי רעיונות מופשטות, כמו האם חשמל היא מלאכת דאורייתא בשבת,  ודווקא קהילה קטנה יכולה לפסוק לגבי נושאים פרטנים, כמו האם הבשר הזה נחשב חֵלב.  אלא שסמכותו של הקהילה הקטנה רק לגבי הקהילה הקטנה.  ולכן היחס בין גודל הקהילה לסמכותה מתחלקת לשתי אופנים, אופן אחת היא הנושאים (גדולה מתעסקת בנושאים מופשטים) ואופן השניה היא סמכות הפסיקה (רק חלה על הקהילה)







Tuesday, April 7, 2020

The Emperor is naked -- deep learning was a good idea but gradient descent is not

Jan 20, 2020

Listening to Eric Weinstein talk about the DISC with his brother Bret made me think about my story.  It is a good conspiracy theory in the making.

Back in the 1990's AI took a turn towards physics, I credit the Hebrew University program started by Daniel Amit and Haim Sompolinsky (both physicists) together with Hana Parnas & A? Abelas (biology) and Tali Tishby (Computer Science -- but really math/physics).  They created the Center for Neural Computation an interdisciplinary center for study of the brain, of which I was a student from its inauguration.

Clearly I am biased, I am more an engineer than a theoretical thinker, and partial differential equations were never my thing.  And I am sure it started innocently enough, a bunch of great minds get together with a solid new idea and apply the tools they have.  Perhaps one of the flags was when Charlie Rosenberg was left out of the circle and forced to find a career elsewhere.  Charlie was more a software engineer than a theoretical mathematician.  Yet, beyond NetTalk which propelled Neural Nets into the public (and created funding opportunities), I credit Charlie for introducing me to the idea that the interdisciplinary study of Neural Networks provides a common language for a diverse set of people to cooperate in ways that would not be possible otherwise.  A very powerful idea that was lost when the language spoken was partial to differential equations.

So here we are many years later and Neural Networks are now taught in university as a gradient descent solution provided by partial differential equations.   Deep learning has taken off and is driven by an industry and university complex that feeds itself.  All this has made NVIDA very happy, they are leading the field in parallel processing neural units (funny thing when you think that my first project with Haim Sompolinsky was programming the ETAN, Intel's neural net chip).  In addition, the professors retain their position of research (and power) as the oracles of partial differential equations, the heart of the deep learning revolution.

However, learning as a field has been narrowed down to minimizing a loss function.  Why does that matter, when everything works so well.

Well, in truth there are some glitches in the armour, Melanie Mitchell just wrote a book on the topic.   Gary Marcus just just held the #AIDebate with Josha Bengio.   But even before that G Hinton had put out there that perhaps the field is gone astray following the Deep learning path.  Frustratingly he then proposed a model that incorporates 'capsules' that are of course trained with gradient descent...

My position has always been that learning is unsupervised, classification is supervised.  This occurred to me when I first met back-propagation in Judith Dayhoff's book and felt drawn to Kohenen's feature maps.  My masters was focused on how to dynamically learn, which led directly to my doctorate thesis.  Not only is learning unsupervised it happens continually, not in batches.

So what has become apparent to me recently is that I can now better describe why Deep Learning with Gradient Descent is broken.

1.  All supervised learning is a method for overfitting and biasing the learning
2.  Even just choosing the training data is a method of weak supervision and hence bias
3.  Embeddings are numeric representations that inshrine these biases into the system

The current approach to deep learning cannot envision a world without a loss function associated with a labeled training set.

Some of this is now changing, 'few shot' and 'one/zero shot' learning is forcing the Deep Learning community to think a little more about what is learning.  Yet they are still falling into the trap of either 'weak' or 'strong' supervision and gradient descent loss functions (everything must be numeric).

The solution requires us to:
1. Correctly define learning (creation of an efficient model -- not memorization, not classification)
2. Correctly measure success of learning (it cannot be a loss function associated with a target!)
3. Leverage symbolic/categorical learning methods to communicate information between levels
4. Develop a multi step learning approach,
4a.  learn independent observations (fully streamable) at multiple scales
4b.  learn dependent observations (weak supervision)
4c.  classify based on  a & b (strong supervision)

Here is my attempt at solving this for ImageNet - https://www.youtube.com/watch?v=HwmqbUVF26g

---
I now understand something else better as well.  I previous wrote:
https://ashlag-cause-and-kook-affect.blogspot.com/2018_07_09_archive.html

which ends with:
Your choice drives the identity of the elements that you will meet. Hence, when I meet a book and you meet the same book, they are not the same, even though they may contain the same content, but because each book is infused with a second identity, through the second-order relationships, they are different books.  Simply said, the words in the book have different meaning to me than to you due to the different social constructs that we live in, each providing a different context to the words and hence a different meaning.

Now I understand that there are two stages, in the first stage I am completely free, and independent, I construct my own personal identity and world view.  In the second stage I limit my field of view to the distribution of friends I have chosen, the reviewer list.  This is the 'weak' supervision coming into play.  I am biasing myself, but within the limited view of the world that has been imposed upon me.

Now I can at anytime (theoretically) move, and find new friends, then my new community will be the element of weak supervision in my life.  The beauty is that I do not need to reshape my personal identity, that can stay fixed all along preserving my core free will (unsupervised development)

-- here is an interesting article pointed out to me by Tuvia Kutscher
Sparse Algorithms are not Stable: A No-free-lunch Theorem, Huan Xu, Constantine Caramanis, Member, IEEE and Shie Mannor, Senior Member, IEEE

I am thinking that intuitively to me their statement is only true for learning algorithms that minimize a global loss function (and are not sensitive to scale).  However, it seems to me that a local (hierarchical) learning methodology can be both sparse and stable.

Thus given a specific scale a local algorithm is stable. (The system may *appear* to be less sparse at a different scale)



















few shot learning...

A metric is the key to learning.  But how to define a metric in a non-numeric space?

The challenge in symbolic learning is measuring the relationship between categories.  The current style in the deep learning community is to convert the categories to numbers and then train models on the numbers.

So the latest and greatest approach in the deep learning community is to create an embedding.  The embedding maps the categories to a numeric space.  Ideally the numeric space is created in a manner that maximizes the space.  Imagine a square where all the data falls into the top right corner, that is not an efficient use of the space.  Similarly the embedding tries to create a space that utilized efficiently, spreading the data in the space.

Well this mean that the embedding creates a space relative to the data set it is trained on.  Hence, the distribution of the training data is critical in determining the space.  When data from a new distribution appears it will not map well into the embedding.

Hence in the few-shot learning world, were very little is available for the training phase and it is assumed new distribution will arrive, it is not enough to just train a classifier, or search in the embedded space, it is necessary to recreate the embedding with the new data, otherwise the classifier will be trapped looking only in the top left corner of the space.

So now I understand what Bengio is saying, you learn the embedding in a broad space, then you use attention to focus only on part of that space for the problem specific part of that space.

Now if the embedding is broad enough, it didn't learn anything.  If the embedding is narrow it learnt the training distribution.  I get that.


Do they do that?

http://papers.nips.cc/paper/6996-prototypical-networks-for-few-shot-learning.pdf


So, it looks like they create the embedding with the entire train dataset?  I could not figure out if the embedding included data that was later defined as validation or test data or only built upon the training data...

—————
I was trying to understand if when you create the embeddings you train the networks on the entire train set? or the entire dataset?

And later when you define a train/test set for the prototypical networks, the train/test is similarly separated.

So for example, if say the embedding is created for categories A,B & C and then you build the prototypical network on A, B & C, then when categories X & Y come along, you utilized the existing embedding network, map X & Y to the embedded space and then import them into the prototypical network?  Or perhaps A,B,C & X,Y are all utilized to create the embedding, but only A,B,C are employed to create the prototypical network?
—————-

what is the analogy in the Ashlag/Kook world?